diff --git a/docs/proposals/omniroute-full-platform-recovery/GOAL-PROMPT.md b/docs/proposals/omniroute-full-platform-recovery/GOAL-PROMPT.md new file mode 100644 index 00000000000..92ad932ba6f --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/GOAL-PROMPT.md @@ -0,0 +1,98 @@ +--- +title: "Goal prompt: deliver OmniRoute full-platform recovery" +lastUpdated: 2026-09-14 +--- + +# Goal prompt + +## Goal + +Deliver [JON-1228](https://linear.app/palermo/issue/JON-1228/omniroute-package-and-accept-a-full-dashboard-memory-fix-candidate) end to end. Build, test, independently accept, deploy and qualify one complete OmniRoute package containing the reviewed memory and stale-pressure fixes while preserving the full dashboard. + +Planning, a plausible patch, a pull request, an API-only build, green health or partial UAT are milestones—not completion. + +## Read first + +- [Overview](./README.md) +- [Product requirements](./PRD.md) +- [Technical specification](./TECHNICAL-SPEC.md) +- [Implementation plan](./IMPLEMENTATION-PLAN.md) +- [UAT plan](./UAT-PLAN.md) +- [Release and rollback](./RELEASE-AND-ROLLBACK.md) +- [Traceability](./TRACEABILITY.md) +- [Risks and decisions](./RISKS-AND-DECISIONS.md) + +## Repository and tracking + +- Repository: `/home/mrburns/Projects/OmniRoute` +- Integration worktree: `/home/mrburns/Projects/OmniRoute/.claude/worktrees/jon-562-563-release-integration` +- Branch: `fix/jon-562-563-release-integration` +- Documentation commit: `4be3a8e7` +- Documentation PR: [#13631](https://github.com/diegosouzapw/OmniRoute/pull/13631) +- Target base: `release/v3.8.51` +- Parent: [JON-564](https://linear.app/palermo/issue/JON-564/omniroute-stop-recurring-memory-outages-and-prove-long-session) +- Delivery: [JON-1228](https://linear.app/palermo/issue/JON-1228/omniroute-package-and-accept-a-full-dashboard-memory-fix-candidate) +- Memory: [JON-562](https://linear.app/palermo/issue/JON-562/omniroute-profile-and-eliminate-long-session-memory-growth), [PR #13623](https://github.com/diegosouzapw/OmniRoute/pull/13623), commit `96fbc8fb` +- Pressure: [JON-563](https://linear.app/palermo/issue/JON-563/omniroute-fix-stale-pressure-admission-lockout), [PR #13618](https://github.com/diegosouzapw/OmniRoute/pull/13618), commit `9af0c0f9` +- Qualification and final UAT: [JON-559](https://linear.app/palermo/issue/JON-559/omniroute-prove-48-hour-stability-and-roll-out-reversibly) +- Containment: [JON-560](https://linear.app/palermo/issue/JON-560/omniroute-contain-memory-and-recover-alive-but-unusable-service) +- Concurrency: [JON-561](https://linear.app/palermo/issue/JON-561/omniroute-calibrate-concurrency-and-bound-queued-memory) +- Upstream memory issue: [#13621](https://github.com/diegosouzapw/OmniRoute/issues/13621) +- Live service: `omniroute-pilot.service`, port `20128` +- User-facing URL: [https://cursor.tail8bb3d0.ts.net:10460/](https://cursor.tail8bb3d0.ts.net:10460/) + +## Current truth + +- The reviewed fixes pass source tests and worked in an API-only package. +- That package was rolled back because it compiled dashboard pages into blank stubs. +- The current full-dashboard installation works visually but contains neither reviewed fix and still grows toward resource-pressure failure. +- The previous `d049af25` qualification is invalid and contributes no elapsed time to the next run. +- API health does not prove the product works. + +## Product boundary + +The deliverable is one package containing the complete dashboard, the full API, the accepted JON-562 memory fix and the accepted JON-563 pressure-recovery fix. + +User-facing release must use the normal full build. `OMNIROUTE_BUILD_BACKEND_ONLY=1`, backend or contributor build profiles, empty page stubs and a second API-only sidecar fail acceptance. + +## Execution + +1. Read the nearest `AGENTS.md`, Linear issues, Graphiti state, checkpoint and handoff. Confirm the exact worktree, branch, lock and live identity. Move JON-1228 to In Progress without weakening its acceptance contract. +2. Reproduce the full-build failure on the exact base and candidate. Capture the first divergence and full import traces. Completion: one deterministic red command proves the actual build defect. +3. Add regression gates that fail when a Client Component imports server-only or Node runtime code, a user-facing release selects backend-only mode, a required dashboard route becomes an empty stub, or API health passes while the browser is blank. +4. Fix the real client/server boundary. Browser modules must remain free of database, browser automation, filesystem, process and other Node-only dependencies. Put server work behind server modules or API routes. Preserve real producer/consumer contracts. Do not hide the failure with aliases, broad externals, ignored build errors or UI stubs. +5. Preserve the independently accepted JON-562 and JON-563 behavior. Any material change to either fix requires its regression and independent review again. +6. Pin every command to Node 24 with `PATH=/usr/local/bin:/usr/bin:/bin` and prove child processes use Node `v24.18.0`. +7. Run focused build-boundary, JON-562 and JON-563 tests; both core typechecks; full unit and Vitest suites; full lint; Prettier on changed files; and `git diff --check`. Reproduce broad failures on the exact base with the same environment before classifying them as inherited. +8. Produce one normal full release using `npm run build:release`. Then pass `npm run check:pack-artifact`, `npm run check:pack-boot` and `npm run check:install-upgrade` against the exact tarball. +9. Prove source SHA, `dist/BUILD_SHA`, package version, tarball SHA-256, installed bundle hashes, configuration digest and running identity all name the same candidate. Reject unsafe archive names, absolute build-host links and missing dashboard assets. +10. Install the exact tarball into an isolated canary with separate data and configuration. Use one full-dashboard process for browser and API acceptance. +11. Run UAT-01 through UAT-09 from [UAT-PLAN.md](./UAT-PLAN.md). Record visible content, final URL, screenshot, browser console errors, failed network requests and reload behavior. A 200 response or screenshot alone does not pass browser UAT. +12. Obtain independent read-only Standards, Spec, release-artifact and UAT-evidence reviews. Reviewers must not edit their reviewed artifact. Repair accepted blockers through the owning coder, then review the changed artifact. +13. After acceptance, commit the exact reviewed diff, push the feature branch, update the upstream PR and attach the PR plus evidence to JON-1228. +14. Before live activation, inspect live requests and socket queues, drain active streams, verify the rollback package and record the current full-dashboard identity. +15. Stage the accepted package in a versioned directory. Atomically activate it, restart `omniroute-pilot.service` and retain the prior full package. +16. Verify the live dashboard: `/login`, `/`, `/dashboard` to `/home`, `/dashboard/logs`, `/dashboard/conversations`, and a first-party `/_next/` JavaScript asset. Verify the same process serves healthy API status, a real small `/v1/messages` request, one bounded representative long request, JON-563 recovery and JON-562 retention behavior. +17. Roll back immediately if any UI or API acceptance row fails. A blank dashboard is an automatic rollback. +18. After live smoke passes, start a new 48-hour JON-559 qualification clock. Record build identity, PIDs, uptime, restart count, RSS, high-water mark, swap, V8 heap, external and ArrayBuffer memory where safely exposed, connections, queues, health/readiness, admission/pressure state and local versus provider errors every five minutes. +19. Repeat dashboard UAT at 0, 24 and 48 hours. Record at least three matched idle checkpoints and idle-to-burst transitions. Any code/configuration change, restart, blank page, unrecovered pressure state, repeated upward retained-memory trend or failed UAT resets or fails the qualification. +20. Keep JON-564, JON-562, JON-563 and JON-559 open until their own acceptance contracts pass. Close JON-1228 only when the exact full package is independently accepted and all evidence is attached. + +## Final report + +State the exact source commit and PR; files changed; full-build root cause; regression and broad test results; package name and SHA-256; review verdicts; UAT-01 through UAT-09 results; deployed `BUILD_SHA` and hashes; dashboard and API evidence; rollback identity; current memory/swap; qualification start, end and status; tickets updated or closed; and every remaining blocker. + +## Definition of done + +- Normal full-dashboard `build:release` passes. +- Pack, boot and upgrade gates pass. +- JON-562 and JON-563 remain present and accepted. +- UAT-01 through UAT-09 pass on one exact package. +- Independent reviews accept code, package and UAT. +- That package is deployed with rollback. +- Live dashboard and API both work. +- A fresh 48-hour qualification is running or completed with immutable evidence. +- Linear contains the PR, evidence and current status. +- No raw heap snapshot, prompt, credential or secret leaves trusted local storage. + +Keep working while any safe, authorised action remains. diff --git a/docs/proposals/omniroute-full-platform-recovery/IMPLEMENTATION-PLAN.md b/docs/proposals/omniroute-full-platform-recovery/IMPLEMENTATION-PLAN.md new file mode 100644 index 00000000000..43ece653a2c --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/IMPLEMENTATION-PLAN.md @@ -0,0 +1,74 @@ +--- +title: "Implementation plan: full-platform recovery" +lastUpdated: 2026-09-14 +--- + +# Implementation plan and Linear ticket drafts + +This plan preserves the JON-564 contract. It adds the missing full-dashboard release path; it does not reopen the reviewed JON-562/JON-563 source changes without evidence. + +## Sequence + +1. **Freeze the recovery candidate.** Start from `d049af25`, inventory its reviewed diff and record the exact full-build failure on the candidate and relevant base. Check: no backend-only environment setting is present in the build invocation. +2. **Repair the full build blocker.** Fix only the source errors that prevent the normal dashboard build. Check: a new regression fails before the repair and passes after; `npm run build:release` exits 0. +3. **Prove the shipped dashboard.** Build/package/boot the same candidate in isolation, then exercise browser UAT. Check: the artifact includes dashboard assets and renders login plus authenticated UI. +4. **Run release-level reliability evidence.** Re-run JON-562/JON-563 focused evidence from the exact package, then begin a clean 48-hour qualification. Check: no candidate changes after the clock starts. +5. **Promote or roll back.** An independent reviewer accepts the claim-bound evidence. The release operator uses a reversible activation and observes both dashboard and API after the change. + +## Ownership and dependencies + +| Work | Owner role | Depends on | Output | +| ---------------------------------------- | ------------------------------------------ | --------------------------------- | ---------------------------------------------------- | +| Full build diagnosis and repair | senior coder; independent read-only review | Candidate and recorded build log | Minimal fix, red/green regression, full-build output | +| Package/dashboard integrity gate | senior coder; independent read-only review | Full build repair | Artifact guard and clean-install proof | +| JON-562/JON-563 integration verification | director coordinates; reviewers judge | Exact full artifact | Source-to-package hash map and focused test results | +| Browser UAT | QA owner with existing test account | Clean boot of exact artifact | UAT record, screenshots, console/network summary | +| 48-hour qualification | release owner; independent QA | UAT PASS, frozen candidate/config | Metrics, error ledger, verdict | +| Activation/rollback | authorised release operator | Independent acceptance | Activation and rollback observation records | + +One writer owns each worktree and file set. Code, tests, release infrastructure, and this documentation must not be edited concurrently in the same files. No deployment, restart, paid provider test, or ticket closure is authorized by this document alone. + +## Linear work item + +Create one high-priority child of JON-564, then replace this placeholder with its ID. JON-559 remains the only ticket for the 48-hour clock, production rollout, rollback, and final UAT. + +### [JON-1228](https://linear.app/palermo/issue/JON-1228/omniroute-package-and-accept-a-full-dashboard-memory-fix-candidate): package and accept a full-dashboard memory-fix candidate + +**Parent:** JON-564. **Blocks:** JON-559. **Blocked by:** JON-562 and JON-563 until their reviewed source changes are in the pinned candidate. **Related evidence:** GitHub [#13621](https://github.com/diegosouzapw/OmniRoute/issues/13621), PR [#13618](https://github.com/diegosouzapw/OmniRoute/pull/13618), and PR [#13623](https://github.com/diegosouzapw/OmniRoute/pull/13623). + +#### Definition of fixed + +PASS: a single pinned `release/v3.8.51` candidate includes accepted changes from PR #13618 and PR #13623, or records a reviewed equivalent for either. + +PASS: `npm run build:release` succeeds with `OMNIROUTE_BUILD_BACKEND_ONLY` absent and `OMNIROUTE_BUILD_PROFILE` not set to `backend` or `contributor`. + +PASS: the candidate tarball is traceable to the full build and `dist/BUILD_SHA` matches its candidate commit. + +PASS: a clean installation of that tarball boots with an isolated data directory and serves the same API and dashboard product. + +PASS: `/login`, Dashboard Home, Request Logs, and Conversations render nonblank content; the first-party JavaScript asset selected from page HTML returns 200 with JavaScript content. + +PASS: an API-only/backend-only build cannot satisfy this ticket’s checks; no backend-only stub marker is present in a dashboard page/route served by the candidate. + +#### Prove it + +From a clean isolated worktree, run `npm run build:release`, `npm run check:pack-artifact`, `npm run check:pack-boot`, and `npm run check:install-upgrade`. Install only the generated tarball with separate data/config paths, then run browser UAT and the agreed no-spend API smoke from that one full-dashboard process. Attach the redacted candidate manifest, hashes, command exits, screenshots, browser console/network summary, and API result. + +#### Coverage + +Add a focused `tests/unit/build/` regression that fails if a release candidate enables backend-only mode or accepts a backend-only dashboard artifact. Preserve the existing package checks and name the actual UAT evidence path before closure. + +## Board repairs and dependencies + +Repair the real Linear parent relation of JON-562, JON-563, JON-559, JON-560, and JON-561 to JON-564 without rewriting their existing acceptance text. The new package ticket is a child of JON-564 and blocks JON-559. JON-561 informs candidate configuration; JON-560 informs canary readiness and blocks JON-559 rather than package assembly. + +## Existing ticket links + +| Ticket | Role in this recovery | +| ------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | +| [JON-564](https://linear.app/palermo/issue/JON-564/omniroute-stop-recurring-memory-outages-and-prove-long-session) | Parent reliability contract; remains open until all acceptance lines pass. | +| [JON-562](https://linear.app/palermo/issue/JON-562/omniroute-profile-and-eliminate-long-session-memory-growth) | Completed-request memory retention fix and profiling evidence. | +| [JON-563](https://linear.app/palermo/issue/JON-563/omniroute-fix-stale-pressure-admission-lockout) | Stale-pressure recovery fix. | +| [JON-559](https://linear.app/palermo/issue/JON-559/omniroute-prove-48-hour-stability-and-roll-out-reversibly) | Qualification and reversible rollout contract. | +| [JON-560](https://linear.app/palermo/issue/JON-560/omniroute-contain-memory-and-recover-alive-but-unusable-service) | Containment/readiness contract. | +| [JON-561](https://linear.app/palermo/issue/JON-561/omniroute-calibrate-concurrency-and-bound-queued-memory) | Measured concurrency and queue-bound policy. | diff --git a/docs/proposals/omniroute-full-platform-recovery/PRD.md b/docs/proposals/omniroute-full-platform-recovery/PRD.md new file mode 100644 index 00000000000..8f05999bc47 --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/PRD.md @@ -0,0 +1,76 @@ +--- +title: "PRD: OmniRoute full-platform recovery" +lastUpdated: 2026-09-14 +--- + +# Product requirements document + +## Problem + +OmniRoute's current full installation again shows the memory-growth risk that JON-562 and JON-563 address. The reviewed integration was briefly deployed as an API-only package. It passed API health and request checks, but users received a blank dashboard because the package had intentionally stubbed the UI at build time. Health alone therefore did not prove a usable product. + +## Product outcome + +Release one exact, traceable candidate that keeps the full OmniRoute dashboard working and contains both reviewed fixes: + +1. JON-562 bounds completed-request detail retention so a large completed request does not retain its backing request body. +2. JON-563 refreshes stale critical resource-pressure state at the real admission path, while retaining rejection for fresh critical pressure. + +## Users and jobs + +| User | Job | Observable outcome | +| ------------------ | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | +| Dashboard operator | Sign in and manage OmniRoute while it is under normal load. | Login and authenticated dashboard routes render usable content and their required client assets load. | +| API client | Submit a long request through `/v1/messages`. | A valid request follows the normal route; a recovered process does not remain stuck returning stale-pressure 503s. | +| Release operator | Promote a candidate safely. | The artifact, source commit, build sentinel, package, running process and recorded evidence match. | +| On-call operator | Recover from a bad candidate. | The prior known-good full-dashboard artifact can be restored with an observed dashboard and API check. | + +## In scope + +- Repair the full-dashboard build blocker that prevents a normal release artifact. +- Integrate the already reviewed JON-562 and JON-563 changes without rewriting their acceptance contract. +- Add artifact and UAT gates that distinguish a live dashboard from an API-only build. +- Qualify the exact full artifact for 48 hours after all code/configuration changes are frozen. +- Produce redacted, claim-bound evidence and a reversible rollout record. + +## Out of scope + +- Replacing OmniRoute, changing provider/model selection, silently truncating context, or reducing required concurrency to make graphs look better. +- Treating the backend-only build profile as a product fix. +- Publishing raw heap snapshots or running provider-spending tests without the existing authority and budget. +- Closing JON-564, JON-562, or JON-563 from unit tests, a health response, or this plan. + +## Requirements + +### Full product surface + +- **FR-1:** A user-facing candidate must be produced by the normal full build path. `OMNIROUTE_BUILD_BACKEND_ONLY` must not be `1`; `OMNIROUTE_BUILD_PROFILE` must not select `backend` or `contributor`. +- **FR-2:** The artifact must serve `/login`, `/`, `/dashboard`, `/home`, `/dashboard/logs`, and `/dashboard/conversations`. `/` is expected to redirect through the application route. +- **FR-3:** A browser must load login, Dashboard Home, Request Logs, and Conversations without a blank document, fatal client error, or failed first-party JavaScript asset. +- **FR-4:** API smoke remains required but is not a substitute for FR-2 or FR-3. `/api/monitoring/health` and `/v1/messages` must remain usable on the same candidate. + +### Reviewed reliability behaviour + +- **FR-5:** After fresh measurements show pressure below recovery thresholds, each supported real admission path recovers within 10 seconds, including after one hour with no admitted work. No unrelated request or restart may be needed. +- **FR-6:** Fresh critical pressure still sheds new work. A telemetry failure may not permanently latch stale critical state or bypass known-fresh critical state. +- **FR-7:** Completed request details have a count, TTL, and byte budget. The JON-562 fixture must show no retained large backing string after drain at the covered 100k, 300k, and 600k-token-equivalent inputs. + +### Qualification and release + +- **FR-8:** A clean 48-hour qualification uses the exact full artifact and frozen configuration. Any candidate code or configuration change resets the clock. +- **FR-9:** The run has no manual or watchdog restart, no sustained local-pressure outage, and no unexplained retained-heap upward trend at matched idle checkpoints. Provider failures are logged separately from local failures. +- **FR-10:** Rollback restores a known-good full-dashboard artifact and proves both dashboard and API use after activation. + +## Success measures + +| Measure | Pass condition | +| ------------ | ---------------------------------------------------------------------------------------------------------------------------------- | +| Build | Full `npm run build:release` exits 0 without a backend-only profile. | +| Package | `dist/BUILD_SHA` matches the candidate commit; `npm run check:pack-artifact` and `npm run check:pack-boot` pass. | +| Dashboard | Login and authenticated dashboard UAT pass with rendered content plus a loaded first-party JavaScript asset. | +| Reliability | JON-562 and JON-563 regressions pass on the packaged candidate, and the 48-hour criteria pass. | +| Traceability | Every pass claim names the ticket, commit, package hash, configuration digest, command, timestamp, and redacted evidence location. | + +## Product acceptance + +All requirements above are PASS/FAIL. A partial pass is **not** a release. In particular, a green API smoke paired with a blank dashboard is FAIL. diff --git a/docs/proposals/omniroute-full-platform-recovery/README.md b/docs/proposals/omniroute-full-platform-recovery/README.md new file mode 100644 index 00000000000..bb3ad2de711 --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/README.md @@ -0,0 +1,42 @@ +--- +title: "OmniRoute full-platform recovery" +lastUpdated: 2026-09-14 +--- + +# OmniRoute full-platform recovery + +This is the recovery delta for [JON-564](https://linear.app/palermo/issue/JON-564/omniroute-stop-recurring-memory-outages-and-prove-long-session). It makes the reviewed JON-562 and JON-563 fixes releasable as a complete OmniRoute product: API and dashboard together. + +## Status + +**Not ready to deploy.** The reviewed integration commit is `d049af25`. Its API-only pilot proved useful API behaviour, but it intentionally compiled the dashboard into empty stubs. Jon observed a blank dashboard, so that artifact was rolled back. The current full installation is restored, but it contains neither reviewed fix. + +The 48-hour monitor for `d049af25` is invalid. The deployed artifact changed before the window completed. Do not carry its elapsed time into a later qualification. + +## Non-negotiable release rule + +Any artifact offered to dashboard users must be built with the dashboard intact. An API-only or backend-only build is allowed only for a clearly labelled API test; it cannot satisfy dashboard UAT, release qualification, or user-facing deployment. + +## Documents + +- [Product requirements](./PRD.md) +- [Technical specification](./TECHNICAL-SPEC.md) +- [Implementation plan and Linear ticket drafts](./IMPLEMENTATION-PLAN.md) +- [User acceptance test plan](./UAT-PLAN.md) +- [Release and rollback runbook](./RELEASE-AND-ROLLBACK.md) +- [Requirement-to-evidence traceability](./TRACEABILITY.md) +- [Risks and decisions](./RISKS-AND-DECISIONS.md) +- [Execution goal prompt](./GOAL-PROMPT.md) + +## Evidence baseline + +| Item | Verified fact | Source | +| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | +| JON-563 | Source-reviewed stale-pressure recovery is in PR [#13618](https://github.com/diegosouzapw/OmniRoute/pull/13618), commit `9af0c0f9`. | PR and Linear ticket | +| JON-562 | Source-reviewed completed-request retention fix is in PR [#13623](https://github.com/diegosouzapw/OmniRoute/pull/13623), commit `96fbc8fb`. | PR and Linear ticket | +| Integration | The two commits are integrated at `d049af25`. | Integration review bundle | +| API pilot | The API pilot returned successful 50k and 100k-token-equivalent `/v1/messages` requests and had zero swap after drain. | JON-562 evidence comment | +| Rejection | That same package was backend-only; UI pages were stubs and the dashboard was blank. | JON-562/JON-563 rollback comments; `scripts/build/backendOnlyPages.mjs` | +| Current blocker | A normal full dashboard build fails on the candidate and exact base; the review recorded 168 Turbopack errors. | Integration review bundle; GitHub [#12732](https://github.com/diegosouzapw/OmniRoute/issues/12732) | + +The evidence bundle and Linear comments are records, not deployment instructions. Do not publish heap snapshots, prompt bodies, headers, cookies, environment values, or credential material. diff --git a/docs/proposals/omniroute-full-platform-recovery/RELEASE-AND-ROLLBACK.md b/docs/proposals/omniroute-full-platform-recovery/RELEASE-AND-ROLLBACK.md new file mode 100644 index 00000000000..f6ad426e21e --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/RELEASE-AND-ROLLBACK.md @@ -0,0 +1,60 @@ +--- +title: "Release and rollback: full OmniRoute recovery" +lastUpdated: 2026-09-14 +--- + +# Release and rollback + +## Release gate + +The release owner may proceed only after all of these are attached to the linked Linear record: + +1. Full `npm run build:release` pass, with no backend-only or contributor profile. +2. Exact artifact package/boot evidence, including matching candidate commit and `dist/BUILD_SHA`. +3. UAT-01 through UAT-09 PASS from [the UAT plan](./UAT-PLAN.md). +4. JON-562 and JON-563 focused evidence on the exact candidate. +5. A 48-hour qualification PASS with frozen candidate/configuration and independent review. +6. A named known-good **full-dashboard** rollback artifact and its identity record. + +The retired `d049af25` API-only package must never be selected as the user-facing rollback target. + +## Prepare without changing live service + +- Record the currently active full artifact identity, process/service identity, package version, build SHA, and redacted configuration digest. +- Verify the rollback artifact has previously passed dashboard UAT, not merely API health. +- Confirm the target release uses isolated state as required by the existing release procedure. Do not share a writable database between candidate and live processes. +- Prepare the evidence directory and ticket links before activation. Record names and digests, never secret values. + +## Controlled activation + +1. Use the project’s existing authorised activation mechanism. Do not replace it with ad hoc global package changes. +2. Activate only the exact approved full-dashboard package. +3. Record the post-activation package/build/process identity. +4. Immediately perform UAT-01, UAT-03, UAT-05, and UAT-06 against the active candidate. These checks prove public login, authenticated UI, a dashboard asset, and API parity. +5. Continue the approved qualification monitoring. It must classify local pressure, process restart, swap/RSS movement, dashboard failure, and provider errors separately. + +## Immediate rollback triggers + +Rollback is required if any of the following occurs: + +- Login or authenticated dashboard becomes blank or unusable. +- A first-party dashboard asset fails, or a new fatal client error prevents normal UI use. +- Local pressure rejects work and does not recover under the JON-563 contract. +- Retained memory grows outside the agreed qualification bound, swap rises persistently, or the candidate requires a manual/watchdog restart. +- API/authentication/streaming regressions affect the approved client journey. +- Candidate identity cannot be proved, or monitoring evidence is missing/corrupted. + +## Rollback procedure + +1. Stop further promotion and record the trigger, timestamp, candidate identity, and safe diagnostic summary. +2. Preserve a redacted evidence bundle. Do not collect raw snapshots or secrets in a rush. +3. Activate the named known-good full-dashboard rollback artifact using the same authorised mechanism. +4. Record the restored artifact/process identity. +5. Re-run UAT-01, UAT-03, UAT-05, and UAT-06 after restoration. A health-only check is insufficient. +6. Update the linked Linear ticket with the failed gate, attached evidence, rollback observation, and next repair action. Keep the parent acceptance contract open. + +## 48-hour qualification rules + +Start only after full UAT passes and the candidate is frozen. Sample at agreed intervals and retain at least three matched idle checkpoints plus recovery transitions. Record RSS, swap, heap/memory fields available without secrets, pressure-observation age, request/queue counters, restart count, error classification, build/package identity, and test workload hash. + +The earlier d049af25 monitor was stopped during rollback. It is historical API evidence only and contributes zero elapsed time to this gate. diff --git a/docs/proposals/omniroute-full-platform-recovery/RISKS-AND-DECISIONS.md b/docs/proposals/omniroute-full-platform-recovery/RISKS-AND-DECISIONS.md new file mode 100644 index 00000000000..27044086f1f --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/RISKS-AND-DECISIONS.md @@ -0,0 +1,37 @@ +--- +title: "Risks and decisions: OmniRoute full-platform recovery" +lastUpdated: 2026-09-14 +--- + +# Risks and decisions + +## Decisions already made + +| Decision | Why | Consequence | +| --------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | +| Reject backend-only output for user-facing deployment. | It deliberately stubs dashboard pages and produced a blank UI. | API-only smoke may continue only as labelled API evidence. | +| Keep JON-562/JON-563 separate from the full-build repair. | The fixes were reviewed independently; the normal build failure is a separate release-base problem. | Do not bury a build repair inside a memory/pressure claim. | +| Restart 48-hour qualification from zero. | The `d049af25` artifact was rolled back before the qualification window completed. | Historical pilot measurements are supportive only. | +| Require dashboard proof after activation and rollback. | API health was green during the blank-dashboard incident. | Health is a component check, not release acceptance. | + +## Active risks + +| Risk | Impact | Control | Decision owner | +| ------------------------------------------------------ | ------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | ------------------- | +| Normal full build remains blocked by the release base. | No safe full artifact can be produced. | Isolate a minimal full-build repair with a failing-before/passing-after regression. | Build owner | +| Build profile accidentally selects backend-only mode. | A superficially healthy API-only package reaches users. | Add build/package guard and record build-mode inputs in evidence. | Build/package owner | +| Artifact identity drifts between test and deployment. | Test results do not prove the running package. | Bind commit, build sentinel, package hash, config digest, and process identity. | Release owner | +| Dashboard assets exist but runtime still blanks. | Static inspection misses a browser-only failure. | Fresh-browser UAT includes content, console, and selected asset check. | QA owner | +| Full test suite has unrelated inherited failures. | A real candidate failure may be misclassified as noise. | Compare exact base and candidate, name each failure, and retain focused mandatory gates. | Director/reviewer | +| Provider calls cost money or introduce upstream noise. | Qualification becomes expensive or ambiguous. | Use deterministic fixtures first; obtain authority/budget before real provider tests. | Jon/release owner | +| Raw profiling data exposes sensitive content. | Security incident during evidence collection. | Allowlisted profiler environment; local analysis; cleanup verification; publish summaries only. | JON-562 owner | +| Rollback restores API but not UI. | The same outage recurs. | Rollback target must be a known-good full-dashboard artifact; execute UAT after restoration. | Release operator | + +## Open decisions + +1. Which minimal repair resolves the current 168-error full dashboard build failure on the recovery candidate? The build ticket must answer from an actual failing log, not from this plan. +2. What representative workload, concurrency, and throughput target will be frozen for the clean 48-hour run? JON-561 owns the measurement decision. +3. What measured RAM/swap limits and readiness/recycle policy protect the host without invalidating required long streams? JON-560 owns this decision. +4. What approval and budget apply to real provider/client qualification after deterministic evidence passes? Jon or the authorised release owner must record it. + +No open decision permits a hidden workaround. If a decision cannot be made from evidence, stop that lane and record the missing authority or measurement. diff --git a/docs/proposals/omniroute-full-platform-recovery/TECHNICAL-SPEC.md b/docs/proposals/omniroute-full-platform-recovery/TECHNICAL-SPEC.md new file mode 100644 index 00000000000..65da1f455fb --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/TECHNICAL-SPEC.md @@ -0,0 +1,72 @@ +--- +title: "Technical specification: full-platform recovery" +lastUpdated: 2026-09-14 +--- + +# Technical specification + +## Existing candidate and failure boundary + +The integration branch contains `ecaa881c` (JON-563) and `d049af25` (JON-562). The reviewed backend-only tarball is retired for user-facing use. It remains evidence for API behaviour only. + +`scripts/build/backendOnlyPages.mjs` defines backend-only mode. In that mode, `scripts/build/build-next-isolated.mjs` calls `stubDashboardPages()`, which replaces App Router UI entry files with null or minimal components during build and restores source afterwards. API route handlers remain, which explains why API health checks passed while the dashboard was blank. + +## Build and package contract + +| Stage | Required input | Required output | Reject when | +| ------------------ | ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | +| Source integration | The two reviewed commits plus the smallest build-blocker repair. | One candidate commit with no unrelated product changes. | The fix changes JON-562/JON-563 behaviour without a new review and regression. | +| Full build | `npm run build:release` with backend-only variables absent. | `.build/next/` intermediates, `dist/`, and `dist/BUILD_SHA`. | The command fails, uses a backend/contributor profile, or writes a mismatched `BUILD_SHA`. | +| Package policy | Built `dist/` and candidate source. | `npm run check:pack-artifact` passes. | Required runtime files are missing, unexpected files leak, or provenance fails. | +| Clean boot | The produced tarball installed into a fresh temporary prefix and separate data directory. | A booted server returning expected health and validation results. | The install falls back to another global package, uses production data, or cannot prove the exact package identity. | +| Browser UAT | The same installed artifact, authenticated test account, and browser. | Saved screenshots, browser-console summary, route/asset evidence. | Login or dashboard is blank, errors fatally, or a first-party asset fails. | + +The deployment candidate must preserve the standard build flow described in `docs/ops/RELEASE_CHECKLIST.md`: `npm run build:release` cleans `.build` and `dist`, builds Next output, assembles the standalone package, then writes the build SHA sentinel. + +## Full-dashboard invariants + +1. The release command runs without `OMNIROUTE_BUILD_BACKEND_ONLY=1` and without `OMNIROUTE_BUILD_PROFILE=backend` or `contributor`. +2. The process that serves `/v1/messages` also serves the dashboard on its configured port. API and UI are one release unit for this recovery. +3. `/` redirects to `/dashboard`, and `/dashboard` redirects to `/home` in the application source. UAT must follow redirects rather than treating either redirect as a blank-page failure. +4. `/login` is a public screen. The authenticated dashboard check must use the project’s existing login flow, not a copied session cookie. +5. A passing `/api/monitoring/health` response proves liveness data only. It cannot pass a dashboard requirement. + +## Artifact proof + +The package ticket must add or extend a deterministic guard at the real build/package seam. It must fail if a user-facing release build selects a backend-only profile or contains `BACKEND_ONLY_STUB_MARKER` in a served dashboard page/route, and pass for a full build. It must prove the shipped output contains both server runtime and dashboard static/client assets; a source-only test of `isBackendOnlyBuild({}) === false` is insufficient. + +The clean-boot check records only: + +- candidate Git commit and `dist/BUILD_SHA`; +- tarball SHA-256 and file count; +- Node path/version and package version; +- redacted configuration digest and fresh data-directory identity; +- HTTP status, response headers needed to classify content, and browser asset status. + +It must not record request bodies, authorization headers, cookie values, environment values, raw heap snapshots, or provider secrets. + +## Runtime behaviour retained from JON-562 and JON-563 + +| Area | Contract | Existing reviewed coverage | +| ------------------------ | ---------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | +| Completed-request memory | Previews are detached and detail retention has an estimated byte budget alongside count/TTL bounds. | `tests/unit/active-request-stream-chunks-lifecycle.test.ts`; `tests/unit/messages-route-memory-profile.test.ts` | +| Long-request profiling | The real `/v1/messages` boundary is profiled with an allowlisted child environment and raw-snapshot cleanup. | `scripts/perf/messages-route-memory-profile.ts` | +| Stale critical pressure | The actual `/v1/messages` admission path triggers a shared bounded refresh; fresh critical pressure still sheds. | `tests/unit/resource-pressure-admission-recovery.test.ts` | +| Pressure runtime | Concurrent refreshes coalesce; a failed/hung sampler has a bounded outcome; disposal releases waiters. | `tests/unit/resource-pressure-admission-recovery.test.ts` | + +## Required test matrix + +| Layer | Test | Baseline failure it catches | Candidate pass evidence | +| ------------- | ----------------------------------------------------------------- | ------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | +| Unit | JON-563 route-level stale-pressure regression | Recovered process returns stale 503 after inactivity. | Actual route admits after fresh below-threshold sample; fresh critical still rejects. | +| Unit/profile | JON-562 detail-retention and 100k/300k/600k profile matrix | Completed detail retains a large backing body. | `largeBackingRetained=false`; cleanup reports no raw snapshots. | +| Build | Full release build, no API-only flags | Full UI cannot compile. | `npm run build:release` exits 0. | +| Package | Artifact policy, boot, and upgrade | A package omits runtime/asset content or boots a different install. | `check:pack-artifact`, `check:pack-boot`, and `check:install-upgrade` pass for the exact tarball. | +| Browser | Login, dashboard, provider/settings navigation, first-party asset | API-only output renders blank UI. | UAT evidence in [UAT plan](./UAT-PLAN.md). | +| Qualification | 48-hour representative run | Long-session memory/recovery regressions return. | Redacted metrics and review verdict satisfy FR-8 and FR-9. | + +## Technical stop conditions + +- Do not build a user-facing package until the normal full build is green. The API-only fallback is rejected. +- Do not start the 48-hour clock until the full artifact, test artifact, configuration digest, and runtime identity match. +- Stop a candidate after two unsuccessful repairs of the same defect. Preserve the failure evidence and obtain a concrete decision before another repair. diff --git a/docs/proposals/omniroute-full-platform-recovery/TRACEABILITY.md b/docs/proposals/omniroute-full-platform-recovery/TRACEABILITY.md new file mode 100644 index 00000000000..2933dda24dd --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/TRACEABILITY.md @@ -0,0 +1,35 @@ +--- +title: "Traceability: OmniRoute full-platform recovery" +lastUpdated: 2026-09-14 +--- + +# Requirement traceability + +| Requirement | Source | Implementation owner | Evidence required | Ticket | +| -------------------------------------------------- | ----------------------------- | -------------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ | +| Full dashboard is mandatory | PRD FR-1 to FR-4 | Build/package owner | Full build command, build-mode record, UAT-01 to UAT-06 | [JON-1228](https://linear.app/palermo/issue/JON-1228/omniroute-package-and-accept-a-full-dashboard-memory-fix-candidate) | +| Candidate includes reviewed memory fix | PRD FR-7 | JON-562 owner | Commit/package hash map; 100k/300k/600k profile evidence | [JON-562](https://linear.app/palermo/issue/JON-562/omniroute-profile-and-eliminate-long-session-memory-growth) | +| Candidate recovers stale pressure safely | PRD FR-5 and FR-6 | JON-563 owner | Route-level red/green test; packaged-candidate result | [JON-563](https://linear.app/palermo/issue/JON-563/omniroute-fix-stale-pressure-admission-lockout) | +| Package preserves dashboard assets | Technical spec artifact proof | Build/package owner | Package check, clean boot, selected asset response | [JON-1228](https://linear.app/palermo/issue/JON-1228/omniroute-package-and-accept-a-full-dashboard-memory-fix-candidate) | +| Dashboard works for a real operator | PRD FR-2 and FR-3 | QA owner | UAT-01 to UAT-05 screenshots, final URLs, console summary | [JON-1228](https://linear.app/palermo/issue/JON-1228/omniroute-package-and-accept-a-full-dashboard-memory-fix-candidate) | +| API and UI are same release unit | PRD FR-4 | QA owner | UAT-06 and UAT-07 identity and request results | [JON-1228](https://linear.app/palermo/issue/JON-1228/omniroute-package-and-accept-a-full-dashboard-memory-fix-candidate) | +| 48-hour evidence is valid | PRD FR-8 and FR-9 | Release/QA owners | Frozen manifest, time series, error ledger, review verdict | [JON-559](https://linear.app/palermo/issue/JON-559/omniroute-prove-48-hour-stability-and-roll-out-reversibly) | +| Unusable service is contained | JON-560 contract | Operations owner | Readiness/fault evidence, controlled drain/recovery record | [JON-560](https://linear.app/palermo/issue/JON-560/omniroute-contain-memory-and-recover-alive-but-unusable-service) | +| Required concurrency is not hidden by a workaround | JON-561 contract | Policy owner | Frozen workload, queue/throughput/memory results | [JON-561](https://linear.app/palermo/issue/JON-561/omniroute-calibrate-concurrency-and-bound-queued-memory) | +| Rollback restores the full product | PRD FR-10 | Release operator | Restored identity plus UAT-01/UAT-03/UAT-06/UAT-07 | [JON-559](https://linear.app/palermo/issue/JON-559/omniroute-prove-48-hour-stability-and-roll-out-reversibly) | + +## Claim-binding template + +Each evidence artifact starts with this record: + +| Field | Required value | +| ------------- | -------------------------------------------------------------- | +| Claim | Exact requirement and UAT/test IDs proved | +| Ticket | Existing JON ID or replacement Linear ID | +| Candidate | Git commit, package version, `dist/BUILD_SHA`, tarball SHA-256 | +| Configuration | Redacted configuration digest and workload hash | +| Execution | UTC and AEST timestamp, command or UAT steps, exit code/result | +| Evidence | Redacted logs, screenshots, measurements, review verdict | +| Limits | What this result does not prove | + +An evidence record without candidate identity, a specific claim, and a result is not acceptance evidence. diff --git a/docs/proposals/omniroute-full-platform-recovery/UAT-PLAN.md b/docs/proposals/omniroute-full-platform-recovery/UAT-PLAN.md new file mode 100644 index 00000000000..994345302b6 --- /dev/null +++ b/docs/proposals/omniroute-full-platform-recovery/UAT-PLAN.md @@ -0,0 +1,49 @@ +--- +title: "UAT plan: full OmniRoute dashboard recovery" +lastUpdated: 2026-09-14 +--- + +# User acceptance test plan + +## Preconditions + +- An authorised QA operator has the exact full-dashboard tarball installed in an isolated candidate environment. +- The candidate’s Git commit, `dist/BUILD_SHA`, tarball SHA-256, package version, Node version, and redacted configuration digest are recorded before testing. +- A pre-existing authorised dashboard test account is available. Do not put credentials, cookies, or tokens in this document or its evidence. +- The candidate is not an API-only build. Verify the build record before opening the browser. + +## Pass/fail rules + +Every row is PASS or FAIL. A redirect expected by the source is not a failure, but a blank page, fatal console error, or failed first-party initial JavaScript asset is a failure. A health response does not pass a UI row. + +| ID | Journey | Steps | Pass evidence | +| ------ | ---------------------- | ------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------ | +| UAT-01 | Public login | Open `/login`; wait for the document and page UI to settle. | Visible login form/content, screenshot, no fatal browser console error. | +| UAT-02 | Root entry | Open `/`; follow redirects. | Application reaches the expected sign-in or dashboard route, not a blank document. | +| UAT-03 | Dashboard Home | Sign in through the normal UI; open `/dashboard` and follow the source redirect to `/home`. | Authenticated dashboard shell and Home content visible; screenshot and final URL recorded. | +| UAT-04 | Request Logs | From the authenticated shell, open `/dashboard/logs`, then reload it. | Request Logs renders usable content before and after reload; screenshot and console summary. | +| UAT-05 | Conversations | From the authenticated shell, open `/dashboard/conversations`, then reload it. | Conversations renders usable content before and after reload; screenshot and console summary. | +| UAT-06 | Static asset | From the UAT-01 or UAT-03 HTML, select one first-party `/_next/` script URL and fetch it. | HTTP 200, JavaScript content type, nonzero response body; URL and headers saved without cookies. | +| UAT-07 | API parity | Call `/api/monitoring/health`; submit the approved no-spend validation request or fixture to `/v1/messages`. | Health and route validation behave as expected on the same candidate identity. | +| UAT-08 | Recovery regression | Use the isolated JON-563 fixture through the actual route. | Stale recovered state admits within 10 seconds; fresh critical state still rejects. | +| UAT-09 | Long-request retention | Run the approved JON-562 100k/300k/600k-token-equivalent profile fixture. | Covered cases report no retained large backing string after drain; raw snapshots are cleaned. | + +## Execution record + +For each run, save a redacted record under the ticket evidence location containing: + +- UAT ID, timestamp, tester, environment label, candidate commit, package SHA-256, and configuration digest; +- final URL, HTTP status, selected asset status/content type/byte count, and browser-console category summary; +- screenshots for UAT-01 through UAT-05; +- test command and exit code for UAT-07 through UAT-09; +- PASS/FAIL and a plain-English failure description. + +Do not attach browser storage, request bodies, provider responses, prompt text, headers that contain credentials, environment files, or heap snapshots. + +## Browser evidence acceptance + +The tester must use a fresh browser context. A screenshot alone is insufficient: pair it with final URL, page content observation, and browser console/network summary. A 200 HTML response alone is insufficient: pair it with a loaded first-party asset and visible content. + +## UAT exit + +UAT passes only when UAT-01 through UAT-09 pass on one exact candidate. Any artifact, source, configuration, or browser-environment change that can affect the claim requires rerunning the affected rows. UAT failure blocks the 48-hour qualification and release. diff --git a/open-sse/config/codebuddyCn.ts b/open-sse/config/codebuddyCn.ts new file mode 100644 index 00000000000..cbc1d6242e1 --- /dev/null +++ b/open-sse/config/codebuddyCn.ts @@ -0,0 +1,3 @@ +// Single source of truth for the CLI/CodeBuddy version used by OAuth, chat and usage. +// Keep this leaf browser-safe because the provider registry is part of the dashboard graph. +export const CODEBUDDY_CN_USER_AGENT = "CLI/2.108.1 CodeBuddy/2.108.1"; diff --git a/open-sse/config/providers/registry/codebuddy-cn/index.ts b/open-sse/config/providers/registry/codebuddy-cn/index.ts index ae951828de0..c18c8b29a85 100644 --- a/open-sse/config/providers/registry/codebuddy-cn/index.ts +++ b/open-sse/config/providers/registry/codebuddy-cn/index.ts @@ -1,4 +1,4 @@ -import { CODEBUDDY_CN_USER_AGENT } from "@/lib/oauth/constants/oauth"; +import { CODEBUDDY_CN_USER_AGENT } from "../../../codebuddyCn.ts"; import type { RegistryEntry } from "../../shared.ts"; /** diff --git a/open-sse/services/model.ts b/open-sse/services/model.ts index ec0080ef838..d4e13bf5e6b 100644 --- a/open-sse/services/model.ts +++ b/open-sse/services/model.ts @@ -1,7 +1,10 @@ import { PROVIDER_ID_TO_ALIAS, PROVIDER_MODELS } from "../config/providerModels.ts"; +import { resolveProviderAlias } from "./providerAlias.ts"; import { resolveWildcardAlias } from "./wildcardRouter.ts"; import { getRegisteredProviderEffortBaseModelId } from "../utils/registeredEffortVariants.ts"; +export { resolveProviderAlias } from "./providerAlias.ts"; + type ProviderModelAliasMap = Record>; type ModelAliasValue = string | { provider?: string; model?: string }; type ModelAliasMap = Record; @@ -27,38 +30,6 @@ export function stripContextWindowSuffix( return modelStr.replace(CONTEXT_WINDOW_SUFFIX_RE, "").trimEnd(); } -// Derive alias→provider mapping from the single source of truth (PROVIDER_ID_TO_ALIAS) -// This prevents the two maps from drifting out of sync -const ALIAS_TO_PROVIDER_ID: Record = {}; -for (const [id, alias] of Object.entries(PROVIDER_ID_TO_ALIAS)) { - if (ALIAS_TO_PROVIDER_ID[alias]) { - console.log( - `[MODEL] Warning: alias "${alias}" maps to both "${ALIAS_TO_PROVIDER_ID[alias]}" and "${id}". Using "${id}".` - ); - } - ALIAS_TO_PROVIDER_ID[alias] = id; -} -// Manual alias overrides — maps slug-style prefixes to canonical provider IDs. -// These live outside the registry because they represent multiple providers -// or backward-compatible slug changes, not a single provider's display name. -// opencode/ → opencode-zen (the main free/open tier; opencode-go is a separate paid tier) -ALIAS_TO_PROVIDER_ID["opencode"] = "opencode-zen"; -// xiaomi/ is the user-visible prefix for MiMo models; register it so -// parseModel("xiaomi/mimo-v2-flash") resolves provider = "xiaomi-mimo" instead -// of falling through to the identity fallback ("xiaomi"). -ALIAS_TO_PROVIDER_ID["xiaomi"] = "xiaomi-mimo"; -// llamacpp/ is the user-visible alias for the llama-cpp self-hosted provider. -// The canonical ID is "llama-cpp" (with a hyphen), but the catalog and user-facing -// prefix is "llamacpp". Register it so parseModel("llamacpp/") resolves -// provider = "llama-cpp" instead of the identity fallback ("llamacpp"). -ALIAS_TO_PROVIDER_ID["llamacpp"] = "llama-cpp"; -// agy/ is the short alias for antigravity provider. -ALIAS_TO_PROVIDER_ID["agy"] = "antigravity"; -// aq/ is the user-visible prefix for the Amazon Q (AWS Builder ID) provider. -// The canonical provider ID is "amazon-q". Register it so parseModel("aq/") -// resolves provider = "amazon-q" instead of falling through to the identity fallback. -ALIAS_TO_PROVIDER_ID["aq"] = "amazon-q"; - // Provider-scoped legacy model aliases. Used to normalize provider/model inputs // and keep backward compatibility when upstream IDs change. const PROVIDER_MODEL_ALIASES: ProviderModelAliasMap = { @@ -123,7 +94,7 @@ const CROSS_PROXY_MODEL_ALIASES_LOWER = Object.fromEntries( // Reverse index: modelId -> providerIds that expose this model const MODEL_TO_PROVIDERS = new Map(); for (const [aliasOrId, models] of Object.entries(PROVIDER_MODELS)) { - const providerId = ALIAS_TO_PROVIDER_ID[aliasOrId] || aliasOrId; + const providerId = resolveProviderAlias(aliasOrId) || aliasOrId; for (const modelEntry of models || []) { const modelId = modelEntry?.id; if (!modelId) continue; @@ -180,31 +151,6 @@ interface ProviderConnectionLike { is_active?: unknown; } -/** - * Resolve provider alias to provider ID - */ -export function resolveProviderAlias(aliasOrId: string | null | undefined): string | null { - if (typeof aliasOrId !== "string") return null; - // Follow the alias chain transitively so intermediate alias-only hops resolve - // to the final target, but STOP as soon as a hop lands on a registered - // provider id (#2901): "oc" must resolve to the no-auth "opencode" provider, - // NOT continue through the manual "opencode" → "opencode-zen" slug override — - // that override is for user-typed `opencode/` prefixes only. Without this - // boundary the no-auth provider becomes unreachable by any prefix. - // Guarded against infinite loops with both a depth limit and a seen-set. - let current = aliasOrId; - const seen = new Set(); - for (let i = 0; i < 10; i++) { - const next = ALIAS_TO_PROVIDER_ID[current]; - if (!next || next === current) return current; - if (next in PROVIDER_ID_TO_ALIAS) return next; - if (seen.has(next)) return next; - seen.add(next); - current = next; - } - return current; -} - /** * #474 — Resolve a bare model name to the selected connection's `defaultModel`. * diff --git a/open-sse/services/providerAlias.ts b/open-sse/services/providerAlias.ts new file mode 100644 index 00000000000..204e98c83d1 --- /dev/null +++ b/open-sse/services/providerAlias.ts @@ -0,0 +1,58 @@ +import { PROVIDER_ID_TO_ALIAS } from "../config/providerModels.ts"; + +// Derive alias→provider mapping from the single source of truth (PROVIDER_ID_TO_ALIAS) +// This prevents the two maps from drifting out of sync +const ALIAS_TO_PROVIDER_ID: Record = {}; +for (const [id, alias] of Object.entries(PROVIDER_ID_TO_ALIAS)) { + if (ALIAS_TO_PROVIDER_ID[alias]) { + console.log( + `[MODEL] Warning: alias "${alias}" maps to both "${ALIAS_TO_PROVIDER_ID[alias]}" and "${id}". Using "${id}".` + ); + } + ALIAS_TO_PROVIDER_ID[alias] = id; +} +// Manual alias overrides — maps slug-style prefixes to canonical provider IDs. +// These live outside the registry because they represent multiple providers +// or backward-compatible slug changes, not a single provider's display name. +// opencode/ → opencode-zen (the main free/open tier; opencode-go is a separate paid tier) +ALIAS_TO_PROVIDER_ID["opencode"] = "opencode-zen"; +// xiaomi/ is the user-visible prefix for MiMo models; register it so +// parseModel("xiaomi/mimo-v2-flash") resolves provider = "xiaomi-mimo" instead +// of falling through to the identity fallback ("xiaomi"). +ALIAS_TO_PROVIDER_ID["xiaomi"] = "xiaomi-mimo"; +// llamacpp/ is the user-visible alias for the llama-cpp self-hosted provider. +// The canonical ID is "llama-cpp" (with a hyphen), but the catalog and user-facing +// prefix is "llamacpp". Register it so parseModel("llamacpp/") resolves +// provider = "llama-cpp" instead of the identity fallback ("llamacpp"). +ALIAS_TO_PROVIDER_ID["llamacpp"] = "llama-cpp"; +// agy/ is the short alias for antigravity provider. +ALIAS_TO_PROVIDER_ID["agy"] = "antigravity"; +// aq/ is the user-visible prefix for the Amazon Q (AWS Builder ID) provider. +// The canonical provider ID is "amazon-q". Register it so parseModel("aq/") +// resolves provider = "amazon-q" instead of falling through to the identity fallback. +ALIAS_TO_PROVIDER_ID["aq"] = "amazon-q"; + +/** + * Resolve provider alias to provider ID + */ +export function resolveProviderAlias(aliasOrId: string | null | undefined): string | null { + if (typeof aliasOrId !== "string") return null; + // Follow the alias chain transitively so intermediate alias-only hops resolve + // to the final target, but STOP as soon as a hop lands on a registered + // provider id (#2901): "oc" must resolve to the no-auth "opencode" provider, + // NOT continue through the manual "opencode" → "opencode-zen" slug override — + // that override is for user-typed `opencode/` prefixes only. Without this + // boundary the no-auth provider becomes unreachable by any prefix. + // Guarded against infinite loops with both a depth limit and a seen-set. + let current = aliasOrId; + const seen = new Set(); + for (let i = 0; i < 10; i++) { + const next = ALIAS_TO_PROVIDER_ID[current]; + if (!next || next === current) return current; + if (next in PROVIDER_ID_TO_ALIAS) return next; + if (seen.has(next)) return next; + seen.add(next); + current = next; + } + return current; +} diff --git a/open-sse/services/usage/codebuddy-cn.ts b/open-sse/services/usage/codebuddy-cn.ts index 5b8837e5351..c81f51eaa5f 100644 --- a/open-sse/services/usage/codebuddy-cn.ts +++ b/open-sse/services/usage/codebuddy-cn.ts @@ -14,7 +14,7 @@ * packs, "Bonus Pack N" for bonus packs (soonest-expiring first). */ -import { CODEBUDDY_CN_USER_AGENT } from "@/lib/oauth/constants/oauth"; +import { CODEBUDDY_CN_USER_AGENT } from "../../config/codebuddyCn.ts"; const USAGE_URL = "https://copilot.tencent.com/v2/billing/meter/get-user-resource"; @@ -83,7 +83,7 @@ function cycleEndMs(acc: TencentAccount): number { function deductionEndMs(acc: TencentAccount): number { const v = acc.DeductionEndTime; - if (typeof v === "number") return (v < 1e12 ? v * 1000 : v); + if (typeof v === "number") return v < 1e12 ? v * 1000 : v; if (typeof v === "string" && /^\d+$/.test(v)) { const n = Number(v); return n < 1e12 ? n * 1000 : n; diff --git a/open-sse/utils/estimateSize.ts b/open-sse/utils/estimateSize.ts index 9eb7f178fd1..84c92fec8b8 100644 --- a/open-sse/utils/estimateSize.ts +++ b/open-sse/utils/estimateSize.ts @@ -79,6 +79,11 @@ function expandContainerFrame(stack: Frame[], frame: Exclude) stack.push({ t: "v", v: (frame.o as Record)[next.value] }); } +export type SizeEstimateResult = + | { status: "complete"; bytes: number } + | { status: "byte-limit"; bytes: number } + | { status: "node-budget"; bytes: number }; + /** * @param byteLimit - early-exit threshold (default ESTIMATE_SIZE_BYTE_LIMIT, * 256 KiB). Pass the actual threshold you're comparing against (see @@ -87,14 +92,18 @@ function expandContainerFrame(stack: Frame[], frame: Exclude) * the byte check and the node-budget fail-closed fallback both key off this * value, not the fixed module constant, when a caller supplies one. */ -export function estimateSizeFast(value: unknown, byteLimit = ESTIMATE_SIZE_BYTE_LIMIT): number { +export function estimateSizeFastResult( + value: unknown, + byteLimit = ESTIMATE_SIZE_BYTE_LIMIT, + nodeBudget = ESTIMATE_SIZE_NODE_BUDGET +): SizeEstimateResult { let bytes = 0; - let visitsLeft = ESTIMATE_SIZE_NODE_BUDGET; + let visitsLeft = nodeBudget; const seen = new WeakSet(); const stack: Frame[] = [{ t: "v", v: value }]; while (stack.length > 0) { - if (visitsLeft <= 0) return byteLimit + 1; + if (visitsLeft <= 0) return { status: "node-budget", bytes }; const frame = stack.pop()!; if (!isValueFrame(frame)) { @@ -109,7 +118,7 @@ export function estimateSizeFast(value: unknown, byteLimit = ESTIMATE_SIZE_BYTE_ const ty = typeof v; if (ty === "string" || ty === "number" || ty === "boolean") { bytes = addPrimitiveBytes(bytes, v as string | number | boolean); - if (bytes > byteLimit) return bytes; + if (bytes > byteLimit) return { status: "byte-limit", bytes }; continue; } if (ty === "object") { @@ -117,7 +126,16 @@ export function estimateSizeFast(value: unknown, byteLimit = ESTIMATE_SIZE_BYTE_ } } - return bytes; + return { status: "complete", bytes }; +} + +export function estimateSizeFast( + value: unknown, + byteLimit = ESTIMATE_SIZE_BYTE_LIMIT, + nodeBudget = ESTIMATE_SIZE_NODE_BUDGET +): number { + const result = estimateSizeFastResult(value, byteLimit, nodeBudget); + return result.status === "node-budget" ? byteLimit + 1 : result.bytes; } export function isSmallEnoughForSemanticCache(value: unknown): boolean { diff --git a/open-sse/utils/resourcePressure.ts b/open-sse/utils/resourcePressure.ts index 57c956182c0..c1f75f9cf59 100644 --- a/open-sse/utils/resourcePressure.ts +++ b/open-sse/utils/resourcePressure.ts @@ -39,6 +39,7 @@ export type ResourcePressureRuntimeOptions = { staleAfterMs?: number; maxStaleMs?: number; retryAfterMs?: number; + admissionRefreshTimeoutMs?: number; samplerDeps?: SampleResourceSignalsDeps; selfRestart?: { enabled?: boolean; @@ -56,6 +57,7 @@ type ResolvedSelfRestart = { }; const SELF_RESTART_DEFAULT_AFTER_MS = 120_000; +const ADMISSION_REFRESH_TIMEOUT_MS = 10_000; function envFlagEnabled(raw: string | undefined): boolean { return raw != null && /^(1|true|yes|on)$/i.test(raw.trim()); @@ -119,6 +121,7 @@ function logCriticalTransitionDiagnostics( export type ResourcePressureRuntime = { check: () => ResourcePressureGuardResult | null; getObservation: () => ResourcePressureObservation; + getAdmissionSeverity: () => Promise; whenRefreshSettled: () => Promise; dispose: () => void; }; @@ -233,6 +236,10 @@ export function createResourcePressureRuntime( const staleAfterMs = requireDuration("staleAfterMs", options.staleAfterMs ?? 1_000); const maxStaleMs = requireDuration("maxStaleMs", options.maxStaleMs ?? 30_000); const retryAfterMs = requireDuration("retryAfterMs", options.retryAfterMs ?? 1_000); + const admissionRefreshTimeoutMs = requireDuration( + "admissionRefreshTimeoutMs", + options.admissionRefreshTimeoutMs ?? ADMISSION_REFRESH_TIMEOUT_MS + ); if (maxStaleMs < staleAfterMs) { throw new RangeError("maxStaleMs must be greater than or equal to staleAfterMs"); } @@ -256,7 +263,18 @@ export function createResourcePressureRuntime( let nextRefreshAtMs = Number.NEGATIVE_INFINITY; let scheduled = false; let inFlight: Promise | null = null; + type RefreshCycle = { + started: boolean; + refreshSettled: Promise; + resolveRefreshSettled: () => void; + admissionReady: Promise; + resolveAdmissionReady: () => void; + admissionTimer: ReturnType | null; + admissionWaiters: number; + }; + let refreshCycle: RefreshCycle | null = null; let disposed = false; + let failOpenLoggedObservation: string | null = null; let criticalSinceMs: number | null = null; let selfRestartFired = false; @@ -300,9 +318,25 @@ export function createResourcePressureRuntime( } }; - const refresh = (): void => { - if (disposed || inFlight) return; + const settleAdmissionCycle = (cycle: RefreshCycle): void => { + if (cycle.admissionTimer) clearTimeout(cycle.admissionTimer); + cycle.admissionTimer = null; + cycle.resolveAdmissionReady(); + }; + + const settleRefreshCycle = (cycle: RefreshCycle): void => { + settleAdmissionCycle(cycle); + cycle.resolveRefreshSettled(); + if (refreshCycle === cycle) refreshCycle = null; + }; + + const refresh = (cycle: RefreshCycle): void => { + if (disposed || inFlight || refreshCycle !== cycle) { + settleRefreshCycle(cycle); + return; + } scheduled = false; + cycle.started = true; inFlight = Promise.resolve() .then(sample) .then((signals) => { @@ -310,6 +344,7 @@ export function createResourcePressureRuntime( const settledAtMs = nowMs(); lastSignals = signals; state = tracker.observe(signals); + failOpenLoggedObservation = null; observeSelfRestart(settledAtMs); lastRefreshAtMs = settledAtMs; nextRefreshAtMs = settledAtMs + staleAfterMs; @@ -319,13 +354,111 @@ export function createResourcePressureRuntime( }) .finally(() => { inFlight = null; + settleRefreshCycle(cycle); }); }; const scheduleRefresh = (): void => { if (disposed || scheduled || inFlight) return; scheduled = true; - schedule(refresh); + let resolveRefreshSettled!: () => void; + let resolveAdmissionReady!: () => void; + const cycle: RefreshCycle = { + started: false, + refreshSettled: new Promise((resolve) => { + resolveRefreshSettled = resolve; + }), + resolveRefreshSettled: () => resolveRefreshSettled(), + admissionReady: new Promise((resolve) => { + resolveAdmissionReady = resolve; + }), + resolveAdmissionReady: () => resolveAdmissionReady(), + admissionTimer: null, + admissionWaiters: 0, + }; + refreshCycle = cycle; + cycle.admissionTimer = setTimeout(() => { + cycle.admissionTimer = null; + cycle.resolveAdmissionReady(); + if (!cycle.started) { + scheduled = false; + nextRefreshAtMs = nowMs() + retryAfterMs; + cycle.resolveRefreshSettled(); + if (refreshCycle === cycle) refreshCycle = null; + } + }, admissionRefreshTimeoutMs); + cycle.admissionTimer.unref?.(); + try { + schedule(() => refresh(cycle)); + } catch { + scheduled = false; + nextRefreshAtMs = nowMs() + retryAfterMs; + settleRefreshCycle(cycle); + } + }; + + const whenRefreshSettled = async (): Promise => { + if (disposed) return; + // The production scheduler intentionally unrefs its Immediate. A caller + // explicitly awaiting freshness must keep this turn alive long enough for + // that scheduled refresh to start. + if (scheduled) await new Promise((resolve) => setImmediate(resolve)); + const pending = refreshCycle?.refreshSettled; + if (pending) await pending; + }; + + const whenAdmissionRefreshReady = async (): Promise => { + const cycle = refreshCycle; + if (!cycle) return; + cycle.admissionWaiters += 1; + cycle.admissionTimer?.ref?.(); + try { + await cycle.admissionReady; + } finally { + cycle.admissionWaiters = Math.max(0, cycle.admissionWaiters - 1); + if (cycle.admissionWaiters === 0) cycle.admissionTimer?.unref?.(); + } + }; + + const cachedSeverity = (): PressureSeverity => { + const cacheAge = lastSignals + ? Math.max(0, nowMs() - lastRefreshAtMs) + : Number.POSITIVE_INFINITY; + if (cacheAge <= maxStaleMs) return state.severity; + if (state.severity === "critical") { + const observationKey = `${state.observedAtMs}|${state.reason}`; + if (failOpenLoggedObservation !== observationKey) { + failOpenLoggedObservation = observationKey; + console.warn( + `[resourcePressure] cached critical observation expired ` + + `(reason=${state.reason} sampleAgeMs=${cacheAge} maxStaleMs=${maxStaleMs}); ` + + "failing open until a fresh sample is available" + ); + } + } + return "normal"; + }; + + const checkImmediateHeap = (): ResourcePressureGuardResult | null => { + let heapUsedMb = 0; + try { + heapUsedMb = immediateHeapUsedMb(); + } catch { + heapUsedMb = 0; + } + const immediate = immediateHeapGuard(heapUsedMb, heapThresholdMb); + if (immediate) { + const now = nowMs(); + state = { + severity: "critical", + reason: "v8_heap_absolute", + elevatedStreak: 0, + recoveryStreak: 0, + lastTransitionAtMs: now, + observedAtMs: now, + }; + } + return immediate; }; // The self-restart circuit measures *sustained* critical time, so it must not @@ -347,26 +480,10 @@ export function createResourcePressureRuntime( return { check() { - let heapUsedMb = 0; - try { - heapUsedMb = immediateHeapUsedMb(); - } catch { - heapUsedMb = 0; - } - const immediate = immediateHeapGuard(heapUsedMb, heapThresholdMb); + const immediate = checkImmediateHeap(); const now = nowMs(); if (now >= nextRefreshAtMs) scheduleRefresh(); - if (immediate) { - state = { - severity: "critical", - reason: "v8_heap_absolute", - elevatedStreak: 0, - recoveryStreak: 0, - lastTransitionAtMs: now, - observedAtMs: now, - }; - return immediate; - } + if (immediate) return immediate; const cacheAge = lastSignals ? Math.max(0, now - lastRefreshAtMs) : Number.POSITIVE_INFINITY; if (cacheAge > maxStaleMs || state.severity !== "critical") { return null; @@ -381,13 +498,28 @@ export function createResourcePressureRuntime( ); }, getObservation: () => ({ signals: lastSignals, state }), - whenRefreshSettled: async () => { - if (scheduled) await new Promise((resolve) => setImmediate(resolve)); - if (inFlight) await inFlight; + async getAdmissionSeverity() { + const staleCritical = state.severity === "critical" && nowMs() >= nextRefreshAtMs; + if (!staleCritical) return this.check() ? "critical" : cachedSeverity(); + + const immediate = checkImmediateHeap(); + if (nowMs() >= nextRefreshAtMs) scheduleRefresh(); + if (immediate) return "critical"; + + // A structural front-door rejection used to return the cached critical + // state before any downstream check could refresh it. Wait only for an + // already-coalesced stale-critical refresh; normal/high traffic stays on + // the non-blocking stale-while-revalidate path. + await whenAdmissionRefreshReady(); + return this.check() ? "critical" : cachedSeverity(); }, + whenRefreshSettled, dispose() { disposed = true; scheduled = false; + const cycle = refreshCycle; + refreshCycle = null; + if (cycle) settleRefreshCycle(cycle); if (selfRestartDriver) { clearInterval(selfRestartDriver); selfRestartDriver = null; @@ -396,23 +528,47 @@ export function createResourcePressureRuntime( }; } -let defaultRuntime = createResourcePressureRuntime(); +const RUNTIME_STORE_KEY = Symbol.for("omniroute.resourcePressure.runtime"); + +type ResourcePressureRuntimeStore = { runtime?: ResourcePressureRuntime }; +type GlobalWithResourcePressureRuntime = typeof globalThis & { + [RUNTIME_STORE_KEY]?: ResourcePressureRuntimeStore; +}; + +function getRuntimeStore(): ResourcePressureRuntimeStore { + const globalWithStore = globalThis as GlobalWithResourcePressureRuntime; + return (globalWithStore[RUNTIME_STORE_KEY] ??= {}); +} + +function getDefaultRuntime(): ResourcePressureRuntime { + const store = getRuntimeStore(); + if (!store.runtime) store.runtime = createResourcePressureRuntime(); + return store.runtime; +} export function checkResourcePressureGuard(): ResourcePressureGuardResult | null { - return defaultRuntime.check(); + return getDefaultRuntime().check(); } export function getResourcePressureObservation(): ResourcePressureObservation { - return defaultRuntime.getObservation(); + return getDefaultRuntime().getObservation(); +} + +/** Freshness-aware pressure decision for structural request admission. */ +export function getAdmissionResourcePressureSeverity(): Promise { + return getDefaultRuntime().getAdmissionSeverity(); } /** Replaces and disposes the process singleton when configuration is reloaded. */ export function reloadResourcePressureRuntime( options: ResourcePressureRuntimeOptions = {} ): ResourcePressureRuntime { - defaultRuntime.dispose(); - defaultRuntime = createResourcePressureRuntime(options); - return defaultRuntime; + const store = getRuntimeStore(); + const replacement = createResourcePressureRuntime(options); + const previous = store.runtime; + store.runtime = replacement; + previous?.dispose(); + return replacement; } export type { diff --git a/package.json b/package.json index ccebb40b095..53872bfcc17 100644 --- a/package.json +++ b/package.json @@ -111,7 +111,7 @@ "build:contributor": "cross-env OMNIROUTE_BUILD_PROFILE=contributor OMNIROUTE_USE_TURBOPACK=0 node scripts/build/build-next-isolated.mjs", "build:cli": "node --import tsx scripts/build/prepublish.ts", "omniroute:verify": "node scripts/check/omniroute-verify.mjs", - "build:release": "rm -rf .build dist && OMNIROUTE_BUILD_SHA=$(git rev-parse --short HEAD) npm run build && npm run build:cli && node scripts/build/write-build-sha.mjs", + "build:release": "node scripts/build/assert-full-release-profile.mjs && rm -rf .build dist && OMNIROUTE_BUILD_SHA=$(git rev-parse --short HEAD) npm run build && npm run build:cli && node scripts/build/write-build-sha.mjs", "build:native:tproxy": "cd src/mitm/tproxy/native && npx --yes node-gyp rebuild", "start": "node scripts/dev/run-next.mjs start", "homolog": "node scripts/homolog/run.mjs", diff --git a/scripts/build/assert-full-release-profile.mjs b/scripts/build/assert-full-release-profile.mjs new file mode 100644 index 00000000000..14dde3bb6a9 --- /dev/null +++ b/scripts/build/assert-full-release-profile.mjs @@ -0,0 +1,10 @@ +#!/usr/bin/env node + +import { isBackendOnlyBuild } from "./backendOnlyPages.mjs"; + +if (isBackendOnlyBuild()) { + console.error( + "[release] Refusing full release: the selected build profile would stub the dashboard." + ); + process.exit(1); +} diff --git a/scripts/build/build-next-isolated.mjs b/scripts/build/build-next-isolated.mjs index 499649f21ba..ec253042bf1 100644 --- a/scripts/build/build-next-isolated.mjs +++ b/scripts/build/build-next-isolated.mjs @@ -311,6 +311,12 @@ export async function main() { const result = await runNextBuild(); const standaloneDir = path.join(distDir, "standalone"); if (result.code === 0 && (await exists(standaloneDir)) && !isContributorBuild()) { + await fs.writeFile( + path.join(standaloneDir, "BUILD_PROFILE"), + isBackendOnlyBuild() ? "backend\n" : "full\n", + "utf8" + ); + try { await fs.cp(path.join(projectRoot, "docs"), path.join(standaloneDir, "docs"), { recursive: true, diff --git a/scripts/build/pack-artifact-policy.ts b/scripts/build/pack-artifact-policy.ts index 378d3a28058..a4093556041 100644 --- a/scripts/build/pack-artifact-policy.ts +++ b/scripts/build/pack-artifact-policy.ts @@ -33,6 +33,7 @@ export const APP_STAGING_REMOVAL_PATHS: string[] = [ export const APP_STAGING_ALLOWED_EXACT_PATHS: string[] = [ ".env.example", + "BUILD_PROFILE", "BUILD_SHA", "docs/openapi.yaml", // #7065: imported by dist/server-ws.mjs; assembleStandalone copies it but without @@ -169,6 +170,7 @@ export const PACK_ARTIFACT_ROOT_ALLOWED_EXACT_PATHS: string[] = [ export const PACK_ARTIFACT_ROOT_ALLOWED_PATH_PREFIXES: string[] = [ "@omniroute/opencode-plugin/", + "@omniroute/opencode-plugin-v2/", "@omniroute/opencode-provider/", "bin/cli/", // Broad open-sse + src source dirs added to package.json "files" in v3.8.21 @@ -189,6 +191,7 @@ export const PACK_ARTIFACT_REQUIRED_PATHS: string[] = [ "dist/src/lib/usage/callLogArtifactWorker.js", "dist/open-sse/vendor/codex-chatgpt-web/adapters/chatgpt-web/mcp-server.js", "dist/open-sse/services/compression/rules/en/filler.json", + "dist/BUILD_PROFILE", "dist/server.js", "dist/server-ws.mjs", "dist/responses-ws-proxy.mjs", @@ -362,6 +365,44 @@ export function findUnexpectedArtifactPaths( .sort(); } +export const REQUIRED_DASHBOARD_ROUTE_PATHS = [ + "dist/.build/next/server/app/login/page.js", + "dist/.build/next/server/app/(dashboard)/home/page.js", + "dist/.build/next/server/app/(dashboard)/dashboard/logs/page.js", + "dist/.build/next/server/app/(dashboard)/dashboard/conversations/page.js", +]; + +export const REQUIRED_DASHBOARD_CLIENT_MANIFEST_PATHS = REQUIRED_DASHBOARD_ROUTE_PATHS.map( + (routePath) => routePath.replace(/page\.js$/, "page_client-reference-manifest.js") +); + +export function findDashboardArtifactProblems( + artifactPaths: string[], + buildProfile: string +): string[] { + const requiredDashboardPaths = [ + ...REQUIRED_DASHBOARD_ROUTE_PATHS, + ...REQUIRED_DASHBOARD_CLIENT_MANIFEST_PATHS, + ]; + const problems = requiredDashboardPaths + .filter((requiredPath) => !artifactPaths.includes(requiredPath)) + .map((requiredPath) => `missing ${requiredPath}`); + + if ( + !artifactPaths.some((artifactPath) => + /^dist\/\.build\/next\/static\/chunks\/.+\.js$/.test(artifactPath) + ) + ) { + problems.push("missing dashboard JavaScript assets under dist/.build/next/static/"); + } + + if (buildProfile.trim() !== "full") { + problems.push(`dist/BUILD_PROFILE must be full, got ${JSON.stringify(buildProfile.trim())}`); + } + + return problems; +} + export function findMissingArtifactPaths( filePaths: string[], requiredPaths: string[] = [] diff --git a/scripts/build/prepublish.ts b/scripts/build/prepublish.ts index 2ea8e6cb89c..d43834fe0dd 100644 --- a/scripts/build/prepublish.ts +++ b/scripts/build/prepublish.ts @@ -27,6 +27,7 @@ import { join, dirname, relative } from "node:path"; import { fileURLToPath } from "node:url"; import { assembleStandalone } from "./assembleStandalone.mjs"; +import { isBackendOnlyBuild } from "./backendOnlyPages.mjs"; import { isNativeExecutable, resolveLocalBinEntry } from "./buildToolRunner.mjs"; import { resolveBundledNpmEntry } from "./resolveNpmEntry.ts"; import { @@ -70,7 +71,8 @@ function runBuildTool( args: readonly string[], options: Parameters[2] ): void { - const localEntry = resolveLocalBinEntry(packageName, binName); + const buildToolRoot = typeof options?.cwd === "string" ? options.cwd : ROOT; + const localEntry = resolveLocalBinEntry(packageName, binName, buildToolRoot); if (localEntry) { if (isNativeExecutable(localEntry)) { execFileSync(localEntry, [...args], options); @@ -175,17 +177,30 @@ if (existsSync(DIST_DIR)) { // .build/next/standalone artifact produced by `npm run build` (build-next-isolated.mjs). // If the artifact is absent we invoke it exactly once. const NEXT_DIST = process.env.NEXT_DIST_DIR || ".build/next"; -const standaloneServerJs = join(ROOT, NEXT_DIST, "standalone", "server.js"); -if (!existsSync(standaloneServerJs)) { - console.log(" 🏗️ .build/next/standalone not found — running `npm run build` once..."); +const standaloneDir = join(ROOT, NEXT_DIST, "standalone"); +const standaloneServerJs = join(standaloneDir, "server.js"); +const standaloneBuildProfile = join(standaloneDir, "BUILD_PROFILE"); +const requiredBuildProfile = isBackendOnlyBuild() ? "backend" : "full"; +const hasMatchingStandalone = + existsSync(standaloneServerJs) && + existsSync(standaloneBuildProfile) && + readFileSync(standaloneBuildProfile, "utf8").trim() === requiredBuildProfile; +if (!hasMatchingStandalone) { + console.log( + ` 🏗️ ${requiredBuildProfile} standalone artifact not found — running \`npm run build\` once...` + ); execFileSync(process.execPath, ["scripts/build/build-next-isolated.mjs"], { cwd: ROOT, stdio: "inherit", }); - if (!existsSync(standaloneServerJs)) { + if ( + !existsSync(standaloneServerJs) || + !existsSync(standaloneBuildProfile) || + readFileSync(standaloneBuildProfile, "utf8").trim() !== requiredBuildProfile + ) { console.error( - "\n ❌ Standalone build not found after `npm run build` at:", - standaloneServerJs + `\n ❌ ${requiredBuildProfile} standalone build not found after \`npm run build\` at:`, + standaloneDir ); console.error(" Make sure next.config.mjs has: output: 'standalone'"); process.exit(1); @@ -495,7 +510,10 @@ if (existsSync(opencodePluginSrc) && existsSync(join(opencodePluginSrc, "package // install never populates its node_modules — and tsup with `dts: true` // needs the plugin's own devDependencies (typescript, @opencode-ai/plugin // types). Without this install a fresh CI publish fails at this step. - if (!existsSync(join(opencodePluginSrc, "node_modules"))) { + const pluginDependenciesReady = + existsSync(join(opencodePluginSrc, "node_modules", "typescript", "package.json")) && + resolveLocalBinEntry("tsup", "tsup", opencodePluginSrc) !== null; + if (!pluginDependenciesReady) { // The plugin's node_modules is gitignored, so a fresh CI checkout // ALWAYS installs here. The registry CDN is intermittently flaky // (onnxruntime-class ETIMEDOUTs to the Microsoft CDN have repeatedly @@ -506,6 +524,7 @@ if (existsSync(opencodePluginSrc) && existsSync(join(opencodePluginSrc, "package const npmEntry = resolveBundledNpmEntry("npm-cli.js"); const installArgs = [ "install", + "--include=dev", "--no-audit", "--no-fund", "--fetch-retries=2", diff --git a/scripts/build/validate-pack-artifact.ts b/scripts/build/validate-pack-artifact.ts index d0d30fbbd6b..9b3f670bfbf 100644 --- a/scripts/build/validate-pack-artifact.ts +++ b/scripts/build/validate-pack-artifact.ts @@ -1,14 +1,10 @@ #!/usr/bin/env node import { execFileSync, spawnSync } from "node:child_process"; -import { existsSync } from "node:fs"; +import { existsSync, readFileSync } from "node:fs"; import { dirname, join } from "node:path"; import { fileURLToPath } from "node:url"; -import { - makeGitAncestryProbe, - readBuildSha, - resolveBuildProvenance, -} from "./buildProvenance.ts"; +import { makeGitAncestryProbe, readBuildSha, resolveBuildProvenance } from "./buildProvenance.ts"; import { MCP_CLOSURE_SPOT_CHECK_PATH, @@ -20,6 +16,7 @@ import { PACK_ARTIFACT_ALLOWED_EXACT_PATHS, PACK_ARTIFACT_ALLOWED_PATH_PREFIXES, PACK_ARTIFACT_REQUIRED_PATHS, + findDashboardArtifactProblems, findMissingArtifactPaths, findUnexpectedArtifactPaths, parseJsonValuesOutput, @@ -149,6 +146,13 @@ try { const missingRequiredPaths: string[] = POLICY_ONLY ? [] : findMissingArtifactPaths(artifactPaths, PACK_ARTIFACT_REQUIRED_PATHS); + const buildProfilePath = join(ROOT, "dist", "BUILD_PROFILE"); + const dashboardProblems: string[] = POLICY_ONLY + ? [] + : findDashboardArtifactProblems( + artifactPaths, + existsSync(buildProfilePath) ? readFileSync(buildProfilePath, "utf8") : "" + ); // #3821 — broad `files` prefixes (open-sse/, src/lib/, ...) would otherwise allow // co-located *.test.* / __tests__ leaks; ban them explicitly on the real pack list. @@ -179,6 +183,13 @@ try { } } + if (dashboardProblems.length > 0) { + console.error("\n❌ The npm publish artifact does not contain a full dashboard:"); + for (const problem of dashboardProblems) { + console.error(` - ${problem}`); + } + } + if (leakedTestPaths.length > 0) { console.error( "\n❌ Test/spec files leaked into the npm publish artifact (tighten package.json files negations):" @@ -203,6 +214,7 @@ try { if ( unexpectedPaths.length > 0 || missingRequiredPaths.length > 0 || + dashboardProblems.length > 0 || leakedTestPaths.length > 0 || missingMcpPaths.length > 0 ) { diff --git a/scripts/perf/messages-route-memory-profile.ts b/scripts/perf/messages-route-memory-profile.ts new file mode 100644 index 00000000000..09dd8f553c8 --- /dev/null +++ b/scripts/perf/messages-route-memory-profile.ts @@ -0,0 +1,990 @@ +/** + * JON-562: bounded memory profile for the real Claude `/v1/messages` request boundary. + * + * The default driver runs each context size in a fresh child process. The worker uses a + * synthetic request and a local fetch stub, so it exercises admission, parsing, translation, + * request logging and SSE cleanup without credentials or provider calls. Raw payloads are never + * written. Artifacts are private (umask 077) and contain only measurements plus V8 profiles of + * the synthetic process. + * + * Usage: + * node --import tsx/esm scripts/perf/messages-route-memory-profile.ts \ + * --output-dir /tmp/omniroute-JON-562 + * + * Defaults: 100k, 300k and 600k token-equivalents; six sequential requests per fresh process; + * concurrency=1; stream=true; cancellation=none. Context size is the only changing factor. + */ +import { spawnSync } from "node:child_process"; +import { createHash } from "node:crypto"; +import fs from "node:fs"; +import inspector from "node:inspector"; +import os from "node:os"; +import path from "node:path"; +import { fileURLToPath, pathToFileURL } from "node:url"; +import v8 from "node:v8"; + +export const BYTES_PER_TOKEN_EQUIVALENT = 4; +export const DEFAULT_TOKEN_EQUIVALENTS = [100_000, 300_000, 600_000] as const; +const PROVENANCE_FILES = [ + "src/lib/usage/completedRequestDetails.ts", + "src/lib/usage/usageHistory.ts", + "tests/unit/active-request-stream-chunks-lifecycle.test.ts", + "scripts/perf/messages-route-memory-profile.ts", + "tests/unit/messages-route-memory-profile.test.ts", +] as const; + +type ClaudePayload = { + model: string; + max_tokens: number; + stream: boolean; + messages: Array<{ role: "user"; content: string }>; +}; + +export type MemoryRow = { + phase: "baseline" | "after_route" | "after_drain" | "settled" | "final"; + elapsedMs: number; + requestIndex?: number; + heapUsedBytes: number; + heapTotalBytes: number; + rssBytes: number; + externalBytes: number; + arrayBuffersBytes: number; + admission: { + activeHeavy: number; + activeHealthyHeadroom: number; + inflightBytes: number; + queuedBytes: number; + waiting: number; + }; +}; + +type GrowthSummary = { + baselineHeapUsedBytes: number; + finalSettledHeapUsedBytes: number; + settledGrowthBytes: number; + settledSlopeBytesPerRequest: number; + growthToWireRatio: number; +}; + +function markerFor(tokenEquivalent: number): string { + return ["JON", "562", tokenEquivalent, "CONTEXT"].join("-"); +} + +/** + * Build the credential-free environment inherited by profiler children. + * @param source - Environment to copy allowlisted runtime fields from. + * @returns A new environment containing only allowlisted fields and fixed test settings. + */ +export function buildWorkerEnv( + source: NodeJS.ProcessEnv | Record = process.env +): NodeJS.ProcessEnv { + const env: NodeJS.ProcessEnv = {}; + for (const key of ["PATH", "HOME", "TMPDIR", "TEMP", "TMP", "LANG", "LC_ALL", "TZ"] as const) { + const value = source[key]; + if (value) env[key] = value; + } + env.NODE_ENV = "test"; + env.APP_LOG_LEVEL = "error"; + return env; +} + +/** + * Build an ASCII-only Claude body with an exact, repeatable serialized size. + * @param tokenEquivalent - Context size expressed at four serialized bytes per token. + * @returns The synthetic body, its marker and exact target wire size. + * @throws {RangeError} If `tokenEquivalent` is invalid or too small for the fixed envelope. + * @throws {Error} If serialization does not match the calculated target size. + */ +export function buildClaudeContextPayload(tokenEquivalent: number): { + body: ClaudePayload; + marker: string; + targetWireBytes: number; +} { + if (!Number.isSafeInteger(tokenEquivalent) || tokenEquivalent < 64) { + throw new RangeError("tokenEquivalent must be an integer >= 64"); + } + const targetWireBytes = tokenEquivalent * BYTES_PER_TOKEN_EQUIVALENT; + const marker = markerFor(tokenEquivalent); + const body: ClaudePayload = { + model: "openai/gpt-4.1", + max_tokens: 8, + stream: true, + messages: [{ role: "user", content: "" }], + }; + const fixedBytes = Buffer.byteLength(JSON.stringify(body), "utf8"); + const contentBytes = targetWireBytes - fixedBytes; + if (contentBytes < marker.length) { + throw new RangeError("tokenEquivalent is too small for the fixed request envelope"); + } + body.messages[0].content = marker + "x".repeat(contentBytes - marker.length); + const actualBytes = Buffer.byteLength(JSON.stringify(body), "utf8"); + if (actualBytes !== targetWireBytes) { + throw new Error(`payload calibration failed: wanted ${targetWireBytes}, got ${actualBytes}`); + } + return { body, marker, targetWireBytes }; +} + +/** + * Calculate post-GC growth from baseline and settled samples only. + * @param rows - Ordered memory samples from one isolated workload. + * @param wireBytes - Exact serialized request size used to calculate the growth ratio. + * @returns Baseline, final, slope and wire-ratio measurements. + * @throws {Error} If the samples contain no baseline or settled row. + */ +export function summarizeSettledGrowth( + rows: Array>, + wireBytes: number +): GrowthSummary { + const baseline = rows.find((row) => row.phase === "baseline"); + const settled = rows.filter((row) => row.phase === "settled"); + if (!baseline || settled.length === 0) { + throw new Error("baseline and settled samples are required"); + } + const final = settled[settled.length - 1]; + const settledGrowthBytes = final.heapUsedBytes - baseline.heapUsedBytes; + const settledSlopeBytesPerRequest = + settled.length < 2 + ? settledGrowthBytes + : (final.heapUsedBytes - settled[0].heapUsedBytes) / (settled.length - 1); + return { + baselineHeapUsedBytes: baseline.heapUsedBytes, + finalSettledHeapUsedBytes: final.heapUsedBytes, + settledGrowthBytes, + settledSlopeBytesPerRequest, + growthToWireRatio: settledGrowthBytes / wireBytes, + }; +} + +function argValue(flag: string): string | undefined { + const index = process.argv.indexOf(flag); + return index >= 0 ? process.argv[index + 1] : undefined; +} + +function positiveIntArg(flag: string, fallback: number): number { + const raw = argValue(flag); + if (raw === undefined) return fallback; + const value = Number(raw); + if (!Number.isSafeInteger(value) || value <= 0) { + throw new RangeError(`${flag} must be a positive integer`); + } + return value; +} + +function privateDirectory(directory: string): void { + fs.mkdirSync(directory, { recursive: true, mode: 0o700 }); + fs.chmodSync(directory, 0o700); +} + +function writePrivateJson(file: string, value: unknown): void { + fs.writeFileSync(file, JSON.stringify(value, null, 2) + "\n", { mode: 0o600 }); + fs.chmodSync(file, 0o600); +} + +function appendPrivateJsonLine(file: string, value: unknown): void { + fs.appendFileSync(file, JSON.stringify(value) + "\n", { mode: 0o600 }); +} + +/** + * Run work while guaranteeing removal of its raw heap-snapshot path. + * @param snapshotFile - Raw snapshot path owned by the operation. + * @param work - Worker/analyzer operation to run before cleanup. + * @returns The fulfilled result from `work`. + * @throws The original work or cleanup error. + */ +export async function withRawSnapshotCleanup( + snapshotFile: string, + work: () => Promise +): Promise { + try { + return await work(); + } finally { + fs.rmSync(snapshotFile, { force: true }); + } +} + +/** + * Remove a worker snapshot unless a successful driver handoff owns it. + * @param snapshotFile - Raw worker snapshot path. + * @param state - Whether the worker completed and the driver accepted ownership. + * @returns Nothing. + */ +export function cleanupWorkerSnapshot( + snapshotFile: string, + state: { workerComplete: boolean; snapshotHandoff: boolean } +): void { + if (!state.workerComplete || !state.snapshotHandoff) { + fs.rmSync(snapshotFile, { force: true }); + } +} + +function forceGc(): void { + if (typeof globalThis.gc !== "function") { + throw new Error("JON-562 worker requires node --expose-gc"); + } + for (let index = 0; index < 4; index += 1) globalThis.gc(); +} + +function delay(ms: number): Promise { + return new Promise((resolve) => setTimeout(resolve, ms)); +} + +function inspectorPost( + session: inspector.Session, + method: string, + params: Record = {} +): Promise { + return new Promise((resolve, reject) => { + session.post(method, params, (error, result) => { + if (error) reject(error); + else resolve(result as T); + }); + }); +} + +async function startAllocationSampling(): Promise<{ + stop: () => Promise>; + disconnect: () => void; +}> { + const session = new inspector.Session(); + session.connect(); + await inspectorPost(session, "HeapProfiler.enable"); + await inspectorPost(session, "HeapProfiler.startSampling", { + samplingInterval: 32 * 1024, + includeObjectsCollectedByMajorGC: true, + includeObjectsCollectedByMinorGC: true, + }); + return { + stop: async () => { + const result = await inspectorPost<{ profile: Record }>( + session, + "HeapProfiler.stopSampling" + ); + return result.profile; + }, + disconnect: () => session.disconnect(), + }; +} + +function openAiSseResponse(): Response { + const encoder = new TextEncoder(); + const frames = [ + `data: ${JSON.stringify({ + id: "chatcmpl_JON562", + object: "chat.completion.chunk", + created: 1, + model: "gpt-4.1", + choices: [{ index: 0, delta: { role: "assistant", content: "ok" } }], + })}\n\n`, + `data: ${JSON.stringify({ + id: "chatcmpl_JON562", + object: "chat.completion.chunk", + created: 1, + model: "gpt-4.1", + choices: [{ index: 0, delta: {}, finish_reason: "stop" }], + usage: { prompt_tokens: 1, completion_tokens: 1, total_tokens: 2 }, + })}\n\n`, + "data: [DONE]\n\n", + ]; + return new Response( + new ReadableStream({ + start(controller) { + for (const frame of frames) controller.enqueue(encoder.encode(frame)); + controller.close(); + }, + }), + { status: 200, headers: { "content-type": "text/event-stream" } } + ); +} + +async function runWorker(): Promise { + process.umask(0o077); + const tokenEquivalent = positiveIntArg("--tokens", 100_000); + const iterations = positiveIntArg("--iterations", 6); + const outputDir = path.resolve(argValue("--output-dir") ?? ""); + if (!argValue("--output-dir")) throw new Error("--output-dir is required in worker mode"); + privateDirectory(outputDir); + + const memoryFile = path.join(outputDir, "memory.jsonl"); + fs.writeFileSync(memoryFile, "", { mode: 0o600 }); + const startedAt = performance.now(); + const rows: MemoryRow[] = []; + let providerCalls = 0; + let harness: Awaited< + ReturnType< + typeof import("../../tests/integration/_chatPipelineHarness.ts").createChatPipelineHarness + > + > | null = null; + let sampler: Awaited> | null = null; + let heapSnapshotFile: string | null = null; + let workerComplete = false; + const snapshotHandoff = process.argv.includes("--snapshot-handoff"); + + try { + const { createChatPipelineHarness } = + await import("../../tests/integration/_chatPipelineHarness.ts"); + harness = await createChatPipelineHarness(`JON-562-${tokenEquivalent}`); + const messagesRoute = await import("../../src/app/api/v1/messages/route.ts"); + const { perConnectionAdmissionController } = + await import("../../src/shared/middleware/chatBodyAdmission.ts"); + const { reloadResourcePressureRuntime } = + await import("../../open-sse/utils/resourcePressure.ts"); + + reloadResourcePressureRuntime({ + heapThresholdMb: null, + immediateHeapUsedMb: () => 1, + sample: async () => ({ + observedAtMs: Date.now(), + v8: { heapUsedBytes: 1, heapLimitBytes: Number.MAX_SAFE_INTEGER }, + process: { + rssBytes: 1, + externalBytes: 0, + arrayBuffersBytes: 0, + availableBytes: null, + constrainedBytes: null, + }, + cgroup: { + currentBytes: null, + maxBytes: null, + highBytes: null, + fileBytes: null, + events: null, + }, + psi: null, + }), + }); + harness.BaseExecutor.RETRY_CONFIG.delayMs = 0; + await harness.resetStorage(); + await harness.seedConnection("openai", { + name: "JON-562-local-stub", + apiKey: "synthetic-not-a-credential", + }); + globalThis.fetch = async () => { + providerCalls += 1; + return openAiSseResponse(); + }; + + const sample = (phase: MemoryRow["phase"], requestIndex?: number): void => { + const usage = process.memoryUsage(); + const admission = perConnectionAdmissionController.snapshot(); + const row: MemoryRow = { + phase, + elapsedMs: Math.round((performance.now() - startedAt) * 100) / 100, + requestIndex, + heapUsedBytes: usage.heapUsed, + heapTotalBytes: usage.heapTotal, + rssBytes: usage.rss, + externalBytes: usage.external, + arrayBuffersBytes: usage.arrayBuffers, + admission: { + activeHeavy: admission.activeHeavy, + activeHealthyHeadroom: admission.activeHealthyHeadroom, + inflightBytes: admission.inflightBytes, + queuedBytes: admission.queuedBytes, + waiting: admission.waiting, + }, + }; + rows.push(row); + appendPrivateJsonLine(memoryFile, row); + }; + + const runRequest = async ( + tokens: number, + requestIndex: number, + measured: boolean + ): Promise => { + const { body, targetWireBytes } = buildClaudeContextPayload(tokens); + const serialized = JSON.stringify(body); + const request = new Request("http://omniroute.invalid/v1/messages", { + method: "POST", + headers: { + "content-type": "application/json", + "content-length": String(targetWireBytes), + accept: "text/event-stream", + }, + body: serialized, + }); + const response = await messagesRoute.POST(request, {}); + if (measured) sample("after_route", requestIndex); + const responseBody = await response.arrayBuffer(); + if (measured) sample("after_drain", requestIndex); + if (response.status !== 200) { + throw new Error( + `route returned ${response.status} (${responseBody.byteLength} response bytes)` + ); + } + return responseBody.byteLength; + }; + + await runRequest(Math.min(tokenEquivalent, 2_048), 0, false); + await delay(50); + forceGc(); + + // Allocation sampling is deliberately completed before the retention time series. The + // inspector profiler retains its own sampled stack records; leaving it enabled would make + // post-GC heap growth look linear even when the request payload itself was collectible. + sampler = await startAllocationSampling(); + await runRequest(tokenEquivalent, 0, false); + const allocationProfile = await sampler.stop(); + const allocationFile = path.join(outputDir, "allocation.heapprofile"); + writePrivateJson(allocationFile, allocationProfile); + sampler.disconnect(); + sampler = null; + + await delay(100); + forceGc(); + sample("baseline"); + + const responseBytes: number[] = []; + for (let requestIndex = 1; requestIndex <= iterations; requestIndex += 1) { + responseBytes.push(await runRequest(tokenEquivalent, requestIndex, true)); + await delay(50); + forceGc(); + sample("settled", requestIndex); + } + + await delay(250); + forceGc(); + sample("final"); + + if (!process.argv.includes("--no-snapshot")) { + forceGc(); + heapSnapshotFile = path.join(outputDir, "post-gc.heapsnapshot"); + v8.writeHeapSnapshot(heapSnapshotFile); + fs.chmodSync(heapSnapshotFile, 0o600); + } + + const { targetWireBytes } = buildClaudeContextPayload(tokenEquivalent); + const growth = summarizeSettledGrowth(rows, targetWireBytes); + const settledRows = rows.filter((row) => row.phase === "settled"); + const released = settledRows.every( + (row) => + row.admission.activeHeavy === 0 && + row.admission.activeHealthyHeadroom === 0 && + row.admission.inflightBytes === 0 && + row.admission.waiting === 0 + ); + if (!released) throw new Error("admission state remained live after an SSE response drained"); + + const manifest = { + ticket: "JON-562", + status: "complete", + route: "/v1/messages", + tokenEquivalent, + bytesPerTokenEquivalent: BYTES_PER_TOKEN_EQUIVALENT, + exactWireBytes: targetWireBytes, + iterations, + providerCalls, + conditions: { + concurrency: 1, + cancellation: "none", + stream: true, + messageCount: 1, + toolCount: 0, + provider: "local-fetch-stub", + networkCalls: 0, + }, + runtime: { + node: process.version, + platform: process.platform, + arch: process.arch, + heapSizeLimitBytes: v8.getHeapStatistics().heap_size_limit, + }, + responseBytes, + allocationSampling: { + intervalBytes: 32 * 1024, + file: path.basename(allocationFile), + requestCount: 1, + measuredWindow: "one post-warmup request before the retention time series", + }, + retentionSeries: { allocationSamplingEnabled: false }, + heapSnapshot: heapSnapshotFile ? path.basename(heapSnapshotFile) : null, + admissionReleasedAfterEveryRequest: released, + growth, + }; + writePrivateJson(path.join(outputDir, "manifest.json"), manifest); + workerComplete = true; + process.stdout.write(JSON.stringify({ outputDir, status: "complete", growth }) + "\n"); + } finally { + try { + if (sampler) { + await sampler.stop().catch(() => undefined); + sampler.disconnect(); + } + await harness?.cleanup(); + } finally { + if (heapSnapshotFile) { + cleanupWorkerSnapshot(heapSnapshotFile, { workerComplete, snapshotHandoff }); + } + } + } +} + +export type HeapSnapshot = { + snapshot: { + meta: { + node_fields: string[]; + node_types: Array; + edge_fields: string[]; + edge_types: Array; + }; + }; + nodes: number[]; + edges: number[]; + strings: string[]; +}; + +function safeLabel(value: string, marker: string): string { + if (value.includes(marker) || value.length > 120) return ""; + if (path.isAbsolute(value)) return `/${path.basename(value)}`; + return value; +} + +/** + * Classify marker-bearing V8 string graphs by their physical backing size. + * @param parsed - Parsed V8 heap snapshot. + * @param marker - Synthetic context marker to locate. + * @param minimumSelfSizeBytes - Minimum backing-graph size treated as retained context. + * @returns Candidate counts, physical sizes, verdict and large component roots. + */ +export function classifyContextBackingCandidates( + parsed: HeapSnapshot, + marker: string, + minimumSelfSizeBytes: number +): { + matchingPreviewNodes: number; + maxSelfSizeBytes: number; + maxBackingRetainedSizeBytes: number; + retainedContextSelfSizeBytes: number; + largeBackingRetained: boolean; + largeNodeIndexes: number[]; +} { + const nodeFields = parsed.snapshot.meta.node_fields; + const nodeWidth = nodeFields.length; + const nodeTypeIndex = nodeFields.indexOf("type"); + const nodeNameIndex = nodeFields.indexOf("name"); + const nodeSelfSizeIndex = nodeFields.indexOf("self_size"); + const edgeCountIndex = nodeFields.indexOf("edge_count"); + const edgeFields = parsed.snapshot.meta.edge_fields; + const edgeWidth = edgeFields.length; + const edgeTypeIndex = edgeFields.indexOf("type"); + const edgeTargetIndex = edgeFields.indexOf("to_node"); + const nodeTypes = parsed.snapshot.meta.node_types[nodeTypeIndex] as string[]; + const edgeTypes = parsed.snapshot.meta.edge_types[edgeTypeIndex] as string[]; + const nodeCount = parsed.nodes.length / nodeWidth; + const matching: Array<{ nodeIndex: number; selfSizeBytes: number }> = []; + const stringParents: Array = new Array(nodeCount); + const stringChildren: Array = new Array(nodeCount); + const isStringNode = (nodeIndex: number): boolean => { + const type = nodeTypes[parsed.nodes[nodeIndex * nodeWidth + nodeTypeIndex]]; + return type === "string" || type === "concatenated string" || type === "sliced string"; + }; + + let edgeOffset = 0; + for (let from = 0; from < nodeCount; from += 1) { + const edgeCount = parsed.nodes[from * nodeWidth + edgeCountIndex]; + for (let local = 0; local < edgeCount; local += 1) { + const type = edgeTypes[parsed.edges[edgeOffset + edgeTypeIndex]]; + const to = parsed.edges[edgeOffset + edgeTargetIndex] / nodeWidth; + if (type === "internal" && isStringNode(from) && isStringNode(to)) { + (stringChildren[from] ??= []).push(to); + (stringParents[to] ??= []).push(from); + } + edgeOffset += edgeWidth; + } + } + + for (let nodeIndex = 0; nodeIndex < nodeCount; nodeIndex += 1) { + const offset = nodeIndex * nodeWidth; + const type = nodeTypes[parsed.nodes[offset + nodeTypeIndex]]; + if (type !== "string" && type !== "concatenated string" && type !== "sliced string") continue; + const name = parsed.strings[parsed.nodes[offset + nodeNameIndex]] ?? ""; + if (name.includes(marker)) { + matching.push({ nodeIndex, selfSizeBytes: parsed.nodes[offset + nodeSelfSizeIndex] }); + } + } + + const componentRoots = new Set(); + for (const entry of matching) { + const queue = [entry.nodeIndex]; + const seen = new Set(); + while (queue.length > 0) { + const node = queue.pop() as number; + if (seen.has(node)) continue; + seen.add(node); + const parents = stringParents[node] ?? []; + if (parents.length === 0) componentRoots.add(node); + else queue.push(...parents); + } + } + + const components = [...componentRoots].map((root) => { + const nodes = new Set(); + const queue = [root]; + let retainedSizeBytes = 0; + while (queue.length > 0) { + const node = queue.pop() as number; + if (nodes.has(node)) continue; + nodes.add(node); + retainedSizeBytes += parsed.nodes[node * nodeWidth + nodeSelfSizeIndex]; + queue.push(...(stringChildren[node] ?? [])); + } + return { root, nodes, retainedSizeBytes }; + }); + const large = components.filter( + (component) => component.retainedSizeBytes >= minimumSelfSizeBytes + ); + const retainedNodes = new Set(); + for (const component of large) { + for (const node of component.nodes) retainedNodes.add(node); + } + return { + matchingPreviewNodes: matching.length, + maxSelfSizeBytes: matching.reduce( + (maximum, entry) => Math.max(maximum, entry.selfSizeBytes), + 0 + ), + maxBackingRetainedSizeBytes: components.reduce( + (maximum, component) => Math.max(maximum, component.retainedSizeBytes), + 0 + ), + retainedContextSelfSizeBytes: [...retainedNodes].reduce( + (total, node) => total + parsed.nodes[node * nodeWidth + nodeSelfSizeIndex], + 0 + ), + largeBackingRetained: large.length > 0, + largeNodeIndexes: large.map((component) => component.root), + }; +} + +function analyzeSnapshot(snapshotFile: string, marker: string, minimumSelfSizeBytes: number) { + const parsed = JSON.parse(fs.readFileSync(snapshotFile, "utf8")) as HeapSnapshot; + const meta = parsed.snapshot.meta; + const nodeFields = meta.node_fields; + const edgeFields = meta.edge_fields; + const nodeWidth = nodeFields.length; + const edgeWidth = edgeFields.length; + const nodeTypeIndex = nodeFields.indexOf("type"); + const nodeNameIndex = nodeFields.indexOf("name"); + const edgeCountIndex = nodeFields.indexOf("edge_count"); + const edgeTypeIndex = edgeFields.indexOf("type"); + const edgeNameIndex = edgeFields.indexOf("name_or_index"); + const edgeTargetIndex = edgeFields.indexOf("to_node"); + const nodeTypes = meta.node_types[nodeTypeIndex] as string[]; + const edgeTypes = meta.edge_types[edgeTypeIndex] as string[]; + const nodeCount = parsed.nodes.length / nodeWidth; + const classification = classifyContextBackingCandidates(parsed, marker, minimumSelfSizeBytes); + const candidateNodes = new Set(classification.largeNodeIndexes); + + const parents: Array | undefined> = + new Array(nodeCount); + let edgeOffset = 0; + for (let from = 0; from < nodeCount; from += 1) { + const nodeOffset = from * nodeWidth; + const edgeCount = parsed.nodes[nodeOffset + edgeCountIndex]; + for (let local = 0; local < edgeCount; local += 1) { + const type = edgeTypes[parsed.edges[edgeOffset + edgeTypeIndex]]; + const rawName = parsed.edges[edgeOffset + edgeNameIndex]; + const to = parsed.edges[edgeOffset + edgeTargetIndex] / nodeWidth; + if (type !== "weak") { + const edgeName = + type === "element" || type === "hidden" ? String(rawName) : parsed.strings[rawName]; + (parents[to] ??= []).push({ from, edgeType: type, edgeName: edgeName ?? "" }); + } + edgeOffset += edgeWidth; + } + } + + const describeNode = (index: number) => { + const offset = index * nodeWidth; + return { + type: nodeTypes[parsed.nodes[offset + nodeTypeIndex]], + name: safeLabel(parsed.strings[parsed.nodes[offset + nodeNameIndex]] ?? "", marker), + }; + }; + + const paths: unknown[] = []; + for (const target of [...candidateNodes].slice(0, 5)) { + const queue: Array<{ node: number; path: Array> }> = [ + { node: target, path: [{ node: describeNode(target) }] }, + ]; + const seen = new Set([target]); + let found: Array> | null = null; + while (queue.length > 0 && !found) { + const current = queue.shift() as { node: number; path: Array> }; + if (current.node === 0 || current.path.length >= 32) { + found = current.path; + break; + } + for (const parent of parents[current.node] ?? []) { + if (seen.has(parent.from)) continue; + seen.add(parent.from); + const nextPath = [ + ...current.path, + { + retainedBy: describeNode(parent.from), + edgeType: parent.edgeType, + edgeName: safeLabel(parent.edgeName, marker), + }, + ]; + if (parent.from === 0) { + found = nextPath; + break; + } + queue.push({ node: parent.from, path: nextPath }); + } + } + paths.push(found ?? [{ node: describeNode(target) }, { finding: "no root within 32 edges" }]); + } + + return { + matchingContextPreviewNodes: classification.matchingPreviewNodes, + maxContextNodeSelfSizeBytes: classification.maxSelfSizeBytes, + maxContextBackingRetainedSizeBytes: classification.maxBackingRetainedSizeBytes, + retainedContextSelfSizeBytes: classification.retainedContextSelfSizeBytes, + minimumLargeBackingSelfSizeBytes: minimumSelfSizeBytes, + largeBackingRetained: classification.largeBackingRetained, + finding: + candidateNodes.size === 0 + ? "Only detached context previews remained; no context-sized backing string crossed the V8 self_size threshold." + : "A context-sized backing string remained after forced GC; redacted root paths follow.", + paths, + }; +} + +/** + * Enforce the physical large-backing retention gate. + * @param result - Analyzer verdict to enforce. + * @returns Nothing. + * @throws {Error} If a context-sized backing string survived forced garbage collection. + */ +export function assertNoLargeBacking(result: { largeBackingRetained: boolean }): void { + if (result.largeBackingRetained) { + throw new Error("context-sized backing string remained after forced GC"); + } +} + +async function runAnalyzer(): Promise { + process.umask(0o077); + const snapshotArg = argValue("--snapshot"); + if (!snapshotArg) throw new Error("analyzer requires --snapshot"); + const snapshotFile = path.resolve(snapshotArg); + const outputArg = argValue("--output-dir"); + const marker = argValue("--marker") ?? ""; + try { + if (!outputArg || !marker) { + throw new Error("analyzer requires --output-dir and --marker"); + } + const minimumSelfSizeBytes = positiveIntArg("--minimum-self-size-bytes", 1_024); + const outputDir = path.resolve(outputArg); + const result = analyzeSnapshot(snapshotFile, marker, minimumSelfSizeBytes); + const snapshotStat = fs.statSync(snapshotFile); + const snapshotSha256 = await sha256File(snapshotFile); + const redactedResult = { + ...result, + snapshot: { + state: "deleted-after-local-analysis", + byteSize: snapshotStat.size, + sha256: snapshotSha256, + }, + }; + writePrivateJson(path.join(outputDir, "retainers.redacted.json"), redactedResult); + assertNoLargeBacking(result); + process.stdout.write(JSON.stringify(redactedResult) + "\n"); + } finally { + fs.rmSync(snapshotFile, { force: true }); + } +} + +function harnessHash(): string { + return createHash("sha256") + .update(fs.readFileSync(new URL(import.meta.url))) + .digest("hex"); +} + +function gitOutput(args: string[]): string { + const result = spawnSync("git", args, { + encoding: "utf8", + env: buildWorkerEnv(process.env), + }); + if (result.status !== 0) { + throw new Error(`git ${args.join(" ")} failed: ${result.stderr.trim()}`); + } + return result.stdout.trim(); +} + +function sha256(value: string | Buffer): string { + return createHash("sha256").update(value).digest("hex"); +} + +function sourceProvenance() { + const testedCommit = gitOutput(["rev-parse", "HEAD"]); + if (!/^[0-9a-f]{40}$/.test(testedCommit)) { + throw new Error(`invalid git HEAD: ${testedCommit || "empty"}`); + } + const branch = gitOutput(["branch", "--show-current"]); + if (!branch) throw new Error("git branch is empty"); + const sourceFiles = Object.fromEntries( + PROVENANCE_FILES.map((file) => { + if (!fs.existsSync(file)) throw new Error(`provenance file missing: ${file}`); + return [file, sha256(fs.readFileSync(file))]; + }) + ); + const trackedDiff = gitOutput(["diff", "--binary", "HEAD", "--", ...PROVENANCE_FILES]); + return { + testedCommit, + branch, + statusPorcelain: gitOutput(["status", "--porcelain=v1", "--untracked-files=all"]) + .split("\n") + .filter(Boolean), + trackedDiffSha256: sha256(trackedDiff), + sourceSetSha256: sha256(JSON.stringify(sourceFiles)), + sourceFiles, + }; +} + +function sha256File(file: string): Promise { + return new Promise((resolve, reject) => { + const hash = createHash("sha256"); + const input = fs.createReadStream(file); + input.on("error", reject); + input.on("data", (chunk) => hash.update(chunk)); + input.on("end", () => resolve(hash.digest("hex"))); + }); +} + +async function runDriver(): Promise { + process.umask(0o077); + const outputDir = path.resolve( + argValue("--output-dir") ?? + path.join(os.tmpdir(), `omniroute-JON-562-${new Date().toISOString().replace(/[:.]/g, "-")}`) + ); + const iterations = positiveIntArg("--iterations", 6); + const tokenCases = (argValue("--tokens") ?? DEFAULT_TOKEN_EQUIVALENTS.join(",")) + .split(",") + .map((raw) => Number(raw)); + if (tokenCases.some((value) => !Number.isSafeInteger(value) || value < 64)) { + throw new RangeError("--tokens must be a comma-separated list of integers >= 64"); + } + privateDirectory(outputDir); + + const cases: unknown[] = []; + const scriptFile = fileURLToPath(import.meta.url); + for (const tokenEquivalent of tokenCases) { + const caseDir = path.join(outputDir, `context-${tokenEquivalent}`); + privateDirectory(caseDir); + const snapshotFile = path.join(caseDir, "post-gc.heapsnapshot"); + await withRawSnapshotCleanup(snapshotFile, async () => { + const worker = spawnSync( + process.execPath, + [ + "--expose-gc", + "--import", + "tsx/esm", + scriptFile, + "--worker", + "--snapshot-handoff", + "--tokens", + String(tokenEquivalent), + "--iterations", + String(iterations), + "--output-dir", + caseDir, + ], + { + cwd: process.cwd(), + encoding: "utf8", + env: buildWorkerEnv(process.env), + timeout: 180_000, + killSignal: "SIGKILL", + } + ); + if (worker.status !== 0) { + throw new Error(`worker ${tokenEquivalent} failed:\n${worker.stdout}\n${worker.stderr}`); + } + + const analyzer = spawnSync( + process.execPath, + [ + "--max-old-space-size=4096", + "--import", + "tsx/esm", + scriptFile, + "--analyze-snapshot", + "--snapshot", + snapshotFile, + "--output-dir", + caseDir, + "--marker", + markerFor(tokenEquivalent), + "--minimum-self-size-bytes", + String(Math.floor((tokenEquivalent * BYTES_PER_TOKEN_EQUIVALENT) / 2)), + ], + { + cwd: process.cwd(), + encoding: "utf8", + env: buildWorkerEnv(process.env), + timeout: 180_000, + killSignal: "SIGKILL", + } + ); + if (analyzer.status !== 0) { + throw new Error( + `snapshot analyzer ${tokenEquivalent} failed:\n${analyzer.stdout}\n${analyzer.stderr}` + ); + } + const retaining = JSON.parse( + fs.readFileSync(path.join(caseDir, "retainers.redacted.json"), "utf8") + ); + assertNoLargeBacking(retaining); + const caseManifestFile = path.join(caseDir, "manifest.json"); + const caseManifest = JSON.parse(fs.readFileSync(caseManifestFile, "utf8")); + caseManifest.heapSnapshot = retaining.snapshot; + writePrivateJson(caseManifestFile, caseManifest); + cases.push(caseManifest); + }); + } + + const provenance = sourceProvenance(); + const workloadManifest = { + ticket: "JON-562", + status: "complete", + testedCommit: provenance.testedCommit, + provenance, + harnessSha256: harnessHash(), + baseBranch: "release/v3.8.51", + inheritedBaseRed: "diegosouzapw/OmniRoute#12732", + variedFactor: "serialized context bytes only", + fixedConditions: { + concurrency: 1, + cancellation: "none", + stream: true, + iterations, + messageCount: 1, + toolCount: 0, + provider: "local-fetch-stub", + externalProviderCalls: 0, + }, + tokenEquivalentCases: tokenCases, + bytesPerTokenEquivalent: BYTES_PER_TOKEN_EQUIVALENT, + byteBudgetEvidence: + "Unit-level only: real-route responses in this checkpoint are small, so the 256-entry cap binds before the 16 MiB byte cap. The route matrix proves sliced backing detachment, not a 16 MiB runtime plateau.", + cases, + }; + writePrivateJson(path.join(outputDir, "workload-manifest.json"), workloadManifest); + process.stdout.write(JSON.stringify({ outputDir, status: "complete" }) + "\n"); +} + +async function main(): Promise { + if (process.argv.includes("--worker")) return runWorker(); + if (process.argv.includes("--analyze-snapshot")) return runAnalyzer(); + return runDriver(); +} + +const invokedPath = process.argv[1] ? pathToFileURL(path.resolve(process.argv[1])).href : ""; +if (invokedPath === import.meta.url) { + main().catch((error: unknown) => { + const message = error instanceof Error ? error.message : String(error); + process.stderr.write(`[JON-562] ${message}\n`); + process.exitCode = 1; + }); +} diff --git a/src/lib/combos/controlCenter.ts b/src/lib/combos/controlCenter.ts index 92a87b2738a..3c16a03e8ab 100644 --- a/src/lib/combos/controlCenter.ts +++ b/src/lib/combos/controlCenter.ts @@ -1,6 +1,6 @@ import { normalizeComboModels, type ComboStep } from "./steps"; import { resolveComboTargetModelStr } from "../../../open-sse/services/combo/opencodeTargetAlias.ts"; -import { resolveProviderAlias } from "../../../open-sse/services/model.ts"; +import { resolveProviderAlias } from "../../../open-sse/services/providerAlias.ts"; type JsonRecord = Record; diff --git a/src/lib/oauth/constants/oauth.ts b/src/lib/oauth/constants/oauth.ts index 8094bb58cd0..5033d0ec3fd 100644 --- a/src/lib/oauth/constants/oauth.ts +++ b/src/lib/oauth/constants/oauth.ts @@ -18,6 +18,7 @@ import { GROK_BUILD_OAUTH_SCOPES, GROK_BUILD_TOKEN_URL, } from "@omniroute/open-sse/config/grokBuild.ts"; +import { CODEBUDDY_CN_USER_AGENT } from "@omniroute/open-sse/config/codebuddyCn.ts"; import { resolvePublicCred } from "@omniroute/open-sse/utils/publicCreds.ts"; import { CURSOR_AGENT_CLI_VERSION } from "@omniroute/open-sse/utils/cursorAgentCliVersion.ts"; import { buildGitLabOAuthEndpoints, GITLAB_DUO_DEFAULT_BASE_URL } from "../gitlab"; @@ -113,7 +114,7 @@ export const QODER_CONFIG = { // (open-sse/services/usage/codebuddy-cn.ts) — a mismatched version string across a // single account's auth vs. chat calls is exactly the kind of internally-inconsistent // client fingerprint Tencent's WAF flags as anomalous (#12702). -export const CODEBUDDY_CN_USER_AGENT = "CLI/2.108.1 CodeBuddy/2.108.1"; +export { CODEBUDDY_CN_USER_AGENT }; export const CODEBUDDY_CN_CONFIG = { baseUrl: "https://copilot.tencent.com", diff --git a/src/lib/usage/callLogArtifactWriter.ts b/src/lib/usage/callLogArtifactWriter.ts index 7cd8452340e..b4b03f2ff7b 100644 --- a/src/lib/usage/callLogArtifactWriter.ts +++ b/src/lib/usage/callLogArtifactWriter.ts @@ -3,9 +3,16 @@ import path from "node:path"; import { fileURLToPath, pathToFileURL } from "node:url"; import { Worker } from "node:worker_threads"; +import { + ESTIMATE_SIZE_NODE_BUDGET, + estimateSizeFastResult, +} from "@omniroute/open-sse/utils/estimateSize.ts"; +import { isAutomatedTestProcess } from "@/shared/utils/testProcess"; import type { CallLogArtifact, CallLogArtifactWriteResult } from "./callLogArtifacts.ts"; +import { parseFileSize } from "../logEnv.ts"; const MAX_QUEUED_JOBS = 128; +export const DEFAULT_MAX_QUEUED_BYTES = 64 * 1024 * 1024; const IDLE_TIMEOUT_MS = 30_000; const CLOSE_TIMEOUT_MS = 2_000; const WARNING_INTERVAL_MS = 30_000; @@ -18,6 +25,7 @@ type WorkerReply = { type QueueItem = { id: number; artifact: CallLogArtifact; + estimatedBytes: number; environment: { pipelineMaxSizeKb?: string; chatDebugFile?: string; @@ -29,6 +37,7 @@ type QueueItem = { let worker: Worker | null = null; let active: QueueItem | null = null; const queue: QueueItem[] = []; +let queuedBytes = 0; let nextId = 1; let idleTimer: NodeJS.Timeout | null = null; let closing = false; @@ -82,6 +91,13 @@ export function resolveCallLogArtifactWorker(context: WorkerResolutionContext = return { workerFile: entryJs ?? cwdJs, execArgv: [] }; } +function getMaxQueuedBytes(): number { + const raw = process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES; + if (!raw) return DEFAULT_MAX_QUEUED_BYTES; + const parsed = parseFileSize(raw); + return parsed > 0 ? parsed : DEFAULT_MAX_QUEUED_BYTES; +} + function clearIdleTimer(): void { if (!idleTimer) return; clearTimeout(idleTimer); @@ -121,6 +137,7 @@ function failOpen(warn = false): void { const failed = active ? [active, ...queue] : [...queue]; active = null; queue.length = 0; + queuedBytes = 0; terminateWorker(); for (const item of failed) item.resolve(null); notifyCloseWaiters(); @@ -138,6 +155,7 @@ function ensureWorker(): Worker { if (!active || reply.id !== active.id) return; const completed = active; active = null; + queuedBytes = Math.max(0, queuedBytes - completed.estimatedBytes); completed.resolve(reply.result); pump(); }); @@ -176,15 +194,29 @@ function pump(): void { export function writeCallArtifactAsync( artifact: CallLogArtifact ): Promise { - if (closing || queue.length >= MAX_QUEUED_JOBS) { + const maxQueuedBytes = getMaxQueuedBytes(); + let estimatedBytes: number; + try { + const estimate = estimateSizeFastResult(artifact, maxQueuedBytes); + estimatedBytes = + estimate.status === "node-budget" + ? Math.max(estimate.bytes, ESTIMATE_SIZE_NODE_BUDGET) + : estimate.bytes; + } catch { + warnRateLimited("[callLogs] Call-log artifact size estimation failed; detail omitted."); + return Promise.resolve(null); + } + if (closing || queue.length >= MAX_QUEUED_JOBS || estimatedBytes > maxQueuedBytes - queuedBytes) { warnRateLimited("[callLogs] Call-log artifact queue unavailable; detail omitted."); return Promise.resolve(null); } + queuedBytes += estimatedBytes; return new Promise((resolve) => { const item = { id: nextId++, artifact, + estimatedBytes, environment: { pipelineMaxSizeKb: process.env.CALL_LOG_PIPELINE_MAX_SIZE_KB, chatDebugFile: process.env.CHAT_DEBUG_FILE, @@ -199,26 +231,49 @@ export function writeCallArtifactAsync( export async function closeCallLogArtifactWriter(timeoutMs = CLOSE_TIMEOUT_MS): Promise { closing = true; - if (!active && queue.length === 0) { - terminateWorker(); - return; - } + try { + if (!active && queue.length === 0) { + terminateWorker(); + return; + } + + if (timeoutMs <= 0) { + failOpen(); + terminateWorker(); + return; + } - if (timeoutMs <= 0) { - failOpen(); + let timeout: NodeJS.Timeout | undefined; + await Promise.race([ + new Promise((resolve) => closeWaiters.push(resolve)), + new Promise((resolve) => { + timeout = setTimeout(resolve, timeoutMs); + timeout.unref?.(); + }), + ]); + if (timeout) clearTimeout(timeout); + if (active || queue.length > 0) failOpen(); terminateWorker(); - return; + } finally { + closing = false; } +} - let timeout: NodeJS.Timeout | undefined; - await Promise.race([ - new Promise((resolve) => closeWaiters.push(resolve)), - new Promise((resolve) => { - timeout = setTimeout(resolve, timeoutMs); - timeout.unref?.(); - }), - ]); - if (timeout) clearTimeout(timeout); - if (active || queue.length > 0) failOpen(); - terminateWorker(); +export function resetCallLogArtifactWriterForTest(): void { + if (!isAutomatedTestProcess()) return; + failOpen(); + closing = false; + lastWarningAt = 0; +} + +export function getCallLogArtifactQueueStats() { + return { + queuedJobs: queue.length, + activeJobs: active ? 1 : 0, + queuedBytes, + maxQueuedJobs: MAX_QUEUED_JOBS, + maxQueuedBytes: getMaxQueuedBytes(), + closing, + hasActiveWorker: Boolean(worker), + }; } diff --git a/src/lib/usage/callLogArtifacts.ts b/src/lib/usage/callLogArtifacts.ts index 1fe14b98e7c..1b2e25ff321 100644 --- a/src/lib/usage/callLogArtifacts.ts +++ b/src/lib/usage/callLogArtifacts.ts @@ -12,9 +12,9 @@ const DATA_DIR = resolveDataDir({ isCloud }); export const CALL_LOGS_DIR = isCloud ? null : path.join(DATA_DIR, "call_logs"); export const MAX_CALL_LOG_ARTIFACT_BYTES = 512 * 1024; -const SIZE_LIMIT_EXCEEDED_REASON = "call_log_artifact_size_limit_exceeded"; -const OMITTED_FOR_SIZE_LIMIT = "[omitted: call log artifact size limit exceeded]"; -const STREAM_CHUNKS_OMITTED_FOR_SIZE_LIMIT = +export const SIZE_LIMIT_EXCEEDED_REASON = "call_log_artifact_size_limit_exceeded"; +export const OMITTED_FOR_SIZE_LIMIT = "[omitted: call log artifact size limit exceeded]"; +export const STREAM_CHUNKS_OMITTED_FOR_SIZE_LIMIT = "[stream chunks omitted: call log artifact size limit exceeded]"; /** @@ -54,7 +54,7 @@ function preserveErrorForSizeLimit(error: unknown): unknown { if (error === null || error === undefined) return null; let serialized: string; try { - serialized = typeof error === "string" ? error : JSON.stringify(error) ?? String(error); + serialized = typeof error === "string" ? error : (JSON.stringify(error) ?? String(error)); } catch { // A circular or unserializable error must not take the whole artifact down. serialized = String(error); diff --git a/src/lib/usage/callLogs.ts b/src/lib/usage/callLogs.ts index 7a2181b0dd6..7f687a04b16 100644 --- a/src/lib/usage/callLogs.ts +++ b/src/lib/usage/callLogs.ts @@ -8,8 +8,10 @@ import fs from "node:fs"; import path from "node:path"; import type { RequestPipelinePayloads } from "@omniroute/open-sse/utils/requestLogger.ts"; +import { estimateSizeFastResult } from "@omniroute/open-sse/utils/estimateSize.ts"; import { sanitizeErrorMessage } from "@omniroute/open-sse/utils/errorSanitization.ts"; import { getDbInstance } from "../db/core"; +import { getCallLogPipelineMaxSizeBytes } from "../logEnv"; import { getRequestDetailLogByCallLogId } from "../db/detailedLogs"; import { shouldPersistToDisk } from "./migrations"; import { getCallLogApiKeyContext } from "./callLogApiKeyContext"; @@ -35,6 +37,10 @@ import { import { pickDisplayValue } from "@/shared/utils/maskEmail"; import { CALL_LOGS_DIR, + MAX_CALL_LOG_ARTIFACT_BYTES, + OMITTED_FOR_SIZE_LIMIT, + SIZE_LIMIT_EXCEEDED_REASON, + STREAM_CHUNKS_OMITTED_FOR_SIZE_LIMIT, readCallArtifact, type CallLogArtifact, type CallLogDetailState, @@ -449,6 +455,43 @@ function getLegacyInlineDetail(id: string) { }; } +function boundPayloadBeforeProtection(value: unknown, maxBytes: number): unknown { + if (value === null || value === undefined) return value; + try { + return estimateSizeFastResult(value, maxBytes).status === "byte-limit" + ? OMITTED_FOR_SIZE_LIMIT + : value; + } catch { + return OMITTED_FOR_SIZE_LIMIT; + } +} + +function boundPipelineBeforeProtection( + value: RequestPipelinePayloads | null, + maxBytes: number +): RequestPipelinePayloads | null | typeof OMITTED_FOR_SIZE_LIMIT { + if (!value) return null; + const streamChunks = value.streamChunks; + if ( + streamChunks && + boundPayloadBeforeProtection(streamChunks, maxBytes) === OMITTED_FOR_SIZE_LIMIT + ) { + value = { + ...value, + streamChunks: { + provider: streamChunks.provider?.length + ? [STREAM_CHUNKS_OMITTED_FOR_SIZE_LIMIT] + : undefined, + openai: streamChunks.openai?.length ? [STREAM_CHUNKS_OMITTED_FOR_SIZE_LIMIT] : undefined, + client: streamChunks.client?.length ? [STREAM_CHUNKS_OMITTED_FOR_SIZE_LIMIT] : undefined, + }, + }; + } + return boundPayloadBeforeProtection(value, maxBytes) === OMITTED_FOR_SIZE_LIMIT + ? OMITTED_FOR_SIZE_LIMIT + : value; +} + async function saveCallLogOperation(entry: any): Promise { try { const apiKeyContext = getCallLogApiKeyContext(); @@ -459,20 +502,46 @@ async function saveCallLogOperation(entry: any): Promise { const apiKeyName = entry.apiKeyName || apiKeyContext?.apiKeyName || null; const noLogEnabled = Boolean(entry.noLog) || (apiKeyId ? isNoLog(apiKeyId) : false); - const protectedRequestBody = noLogEnabled ? null : protectPayloadForLog(entry.requestBody); + const rawPipelinePayloads = entry.pipelinePayloads ?? entry.pipeline ?? null; + const maxArtifactBytes = rawPipelinePayloads + ? getCallLogPipelineMaxSizeBytes() + : MAX_CALL_LOG_ARTIFACT_BYTES; + const boundedRequestBody = noLogEnabled + ? null + : boundPayloadBeforeProtection(entry.requestBody, maxArtifactBytes); + const boundedResponseBody = noLogEnabled + ? null + : boundPayloadBeforeProtection(entry.responseBody, maxArtifactBytes); + const boundedPipelinePayloads = noLogEnabled + ? null + : boundPipelineBeforeProtection(rawPipelinePayloads, maxArtifactBytes); + const protectedRequestBody = noLogEnabled + ? null + : boundedRequestBody === OMITTED_FOR_SIZE_LIMIT + ? OMITTED_FOR_SIZE_LIMIT + : protectPayloadForLog(boundedRequestBody); const responseStatus = Number(entry.status); const failedResponse = Number.isFinite(responseStatus) && responseStatus >= 400; const protectedResponseBody = noLogEnabled ? null - : failedResponse - ? protectErrorPayloadForLog(entry.responseBody) - : protectPayloadForLog(entry.responseBody); + : boundedResponseBody === OMITTED_FOR_SIZE_LIMIT + ? OMITTED_FOR_SIZE_LIMIT + : failedResponse + ? protectErrorPayloadForLog(boundedResponseBody) + : protectPayloadForLog(boundedResponseBody); const protectedPipelinePayloads = noLogEnabled ? null - : protectPipelinePayloads( - entry.pipelinePayloads ?? entry.pipeline ?? null, - failedResponse ? responseStatus : undefined - ); + : boundedPipelinePayloads === OMITTED_FOR_SIZE_LIMIT + ? { + error: { + _omniroute_truncated: true, + reason: SIZE_LIMIT_EXCEEDED_REASON, + }, + } + : protectPipelinePayloads( + boundedPipelinePayloads, + failedResponse ? responseStatus : undefined + ); const protectedError = sanitizeErrorForLog(entry.error); // Bridges the window before this row's own artifact write (queued below, diff --git a/src/lib/usage/completedRequestDetails.ts b/src/lib/usage/completedRequestDetails.ts index ac91b736e20..03d1613e7eb 100644 --- a/src/lib/usage/completedRequestDetails.ts +++ b/src/lib/usage/completedRequestDetails.ts @@ -3,12 +3,44 @@ import type { PendingRequestDetail } from "./usageHistory"; const COMPLETED_DETAIL_TTL_MS = 120_000; const MAX_COMPLETED_DETAILS = 256; +/** + * JON-562: completed details are a short-lived dashboard bridge, not a second payload store. + * The 16 MiB estimated cache payload budget keeps room for normal bridge entries while bounding + * the strings and object fields this module accounts for. It is not a process-memory ceiling. + */ +export const MAX_COMPLETED_DETAILS_BYTES = 16 * 1024 * 1024; const completedDetails = new Map(); const completedDetailTimers = new Map>(); +const completedDetailBytes = new Map(); +let totalCompletedDetailBytes = 0; + +function estimateRetainedBytes(value: unknown, seen = new WeakSet()): number { + if (value === null || value === undefined) return 0; + if (typeof value === "string") return Buffer.byteLength(value, "utf8"); + if (typeof value === "number" || typeof value === "bigint") return 8; + if (typeof value === "boolean") return 4; + if (typeof value !== "object" || seen.has(value)) return 0; + seen.add(value); + + if (Array.isArray(value)) { + return 32 + value.reduce((total, entry) => total + estimateRetainedBytes(entry, seen), 0); + } + + let bytes = 64; + for (const [key, entry] of Object.entries(value as Record)) { + bytes += Buffer.byteLength(key, "utf8") + estimateRetainedBytes(entry, seen); + } + return bytes; +} function deleteCompletedDetail(id: string) { completedDetails.delete(id); + totalCompletedDetailBytes = Math.max( + 0, + totalCompletedDetailBytes - (completedDetailBytes.get(id) ?? 0) + ); + completedDetailBytes.delete(id); const existingTimer = completedDetailTimers.get(id); if (existingTimer) { clearTimeout(existingTimer); @@ -17,7 +49,10 @@ function deleteCompletedDetail(id: string) { } function trimCompletedDetails() { - while (completedDetails.size > MAX_COMPLETED_DETAILS) { + while ( + completedDetails.size > MAX_COMPLETED_DETAILS || + totalCompletedDetailBytes > MAX_COMPLETED_DETAILS_BYTES + ) { const oldestId = completedDetails.keys().next().value; if (!oldestId) break; deleteCompletedDetail(oldestId); @@ -28,17 +63,61 @@ export function getCompletedDetails(): Map { return completedDetails; } -export function storeCompletedDetail(detail: PendingRequestDetail) { - completedDetails.set(detail.id, detail); +/** + * Read the estimated payload bytes currently accounted to the completed-detail cache. + * @returns The cache's estimated payload-byte total. + */ +export function getCompletedDetailsByteSize(): number { + return totalCompletedDetailBytes; +} + +/** + * Read the completed-detail cache counters. + * @returns Entry, cleanup-timer and estimated payload-byte counts. + */ +export function getCompletedDetailsCacheStats(): { + entries: number; + cleanupTimers: number; + bytes: number; +} { + return { + entries: completedDetails.size, + cleanupTimers: completedDetailTimers.size, + bytes: totalCompletedDetailBytes, + }; +} + +/** + * Store a detached completed-request preview. + * @param detail - Completed request detail to detach and cache. + * @returns `true` only when the entry remains cached after count and byte-budget eviction. + * @throws If `detail` contains a value that `structuredClone` cannot copy. + */ +export function storeCompletedDetail(detail: PendingRequestDetail): boolean { + const inputBytes = estimateRetainedBytes(detail); + if (inputBytes > MAX_COMPLETED_DETAILS_BYTES) { + deleteCompletedDetail(detail.id); + return false; + } + + // `truncatePendingPreview()` uses String#slice. V8 may represent that short preview as a + // sliced string whose hidden parent is the full multi-megabyte request. A structured clone + // materializes the visible preview into cache-owned storage and drops the pending graph. + const detached = structuredClone(detail); + const detachedBytes = estimateRetainedBytes(detached); + totalCompletedDetailBytes -= completedDetailBytes.get(detail.id) ?? 0; + completedDetails.set(detail.id, detached); + completedDetailBytes.set(detail.id, detachedBytes); + totalCompletedDetailBytes += detachedBytes; trimCompletedDetails(); + return completedDetails.has(detail.id); } export function scheduleCompletedDetailCleanup(id: string) { const existingTimer = completedDetailTimers.get(id); if (existingTimer) clearTimeout(existingTimer); const timer = setTimeout(() => { - completedDetails.delete(id); - completedDetailTimers.delete(id); + deleteCompletedDetail(id); }, COMPLETED_DETAIL_TTL_MS); timer.unref?.(); completedDetailTimers.set(id, timer); @@ -48,6 +127,8 @@ export function clearCompletedDetails() { for (const timer of completedDetailTimers.values()) clearTimeout(timer); completedDetailTimers.clear(); completedDetails.clear(); + completedDetailBytes.clear(); + totalCompletedDetailBytes = 0; } function isUnset(value: unknown): boolean { diff --git a/src/lib/usage/usageHistory.ts b/src/lib/usage/usageHistory.ts index a6fb9a7d3af..82743d1272a 100644 --- a/src/lib/usage/usageHistory.ts +++ b/src/lib/usage/usageHistory.ts @@ -465,9 +465,11 @@ function finalizePendingDetailAt( completedAt, durationMs: Math.max(0, completedAt - details[index].startedAt), }; - storeCompletedDetail(updated); - maybeEnrichCompletedDetail(updated, connectionId); - scheduleCompletedDetailCleanup(updated.id); + const storedCompletedDetail = storeCompletedDetail(updated); + if (storedCompletedDetail) { + maybeEnrichCompletedDetail(updated, connectionId); + scheduleCompletedDetailCleanup(updated.id); + } details.splice(index, 1); pendingById.delete(updated.id); diff --git a/src/shared/middleware/chatBodyAdmission.ts b/src/shared/middleware/chatBodyAdmission.ts index f838e92224c..dbdcaa91872 100644 --- a/src/shared/middleware/chatBodyAdmission.ts +++ b/src/shared/middleware/chatBodyAdmission.ts @@ -38,6 +38,7 @@ import { type IngestBudgetAcquireResult, } from "./ingestByteAdmission"; import { + getAdmissionResourcePressureSeverity, getResourcePressureObservation, type PressureSeverity, } from "@omniroute/open-sse/utils/resourcePressure.ts"; @@ -217,8 +218,8 @@ export type ChatAdmissionShedReason = | "inflight_bytes_budget" | "resource_pressure"; -/** Read cached pressure severity; sampling failures must not cause false sheds. */ -export function defaultPressureSeverity(): PressureSeverity { +/** Cached pressure for diagnostics only; request decisions use the freshness-aware async seam. */ +function cachedPressureSeverityForDiagnostics(): PressureSeverity { try { return getResourcePressureObservation().state.severity; } catch { @@ -773,7 +774,7 @@ export const perConnectionAdmissionController = new PerConnectionAdmissionContro budget: { maxInflightBytes: productionIngestBudget.bytes, budgetSource: productionIngestBudget.source, - checkPressureSeverity: defaultPressureSeverity, + checkPressureSeverity: cachedPressureSeverityForDiagnostics, }, } ); @@ -1005,7 +1006,15 @@ export async function admitChatRequest( // #503-fanout: shed before spending any bytes on ingestion when the process // is under genuine critical resource pressure. No-op for every controller a // test constructs directly (default severity is always "normal"). - if (controller.pressureSeverity() === "critical") { + let pressureSeverity: PressureSeverity; + try { + pressureSeverity = options.controller + ? controller.pressureSeverity() + : await getAdmissionResourcePressureSeverity(); + } catch { + pressureSeverity = "normal"; + } + if (pressureSeverity === "critical") { controller.recordShed("resource_pressure", sessionId); return { admit: false, response: resourcePressureRejectionResponse() }; } @@ -1039,9 +1048,8 @@ export async function admitChatRequest( // every controller a test constructs directly, so this resolves // synchronously true there — only the production singleton (built with a // real host-derived budget) is ever actually gated by it. - const severity = controller.pressureSeverity(); const budgetWaitMs = - severity === "high" ? queueMs : Math.min(queueMs, INGEST_NORMAL_MAX_WAIT_MS); + pressureSeverity === "high" ? queueMs : Math.min(queueMs, INGEST_NORMAL_MAX_WAIT_MS); const budgetResult = await controller.acquireBudgetWithin( bytes, budgetWaitMs, diff --git a/tests/unit/active-request-stream-chunks-lifecycle.test.ts b/tests/unit/active-request-stream-chunks-lifecycle.test.ts index 35fd1cbc2d8..6ddd3950629 100644 --- a/tests/unit/active-request-stream-chunks-lifecycle.test.ts +++ b/tests/unit/active-request-stream-chunks-lifecycle.test.ts @@ -10,6 +10,7 @@ process.env.DATA_DIR = TEST_DATA_DIR; const core = await import("../../src/lib/db/core.ts"); const usageHistory = await import("../../src/lib/usage/usageHistory.ts"); const callLogs = await import("../../src/lib/usage/callLogs.ts"); +const completedRequestDetails = await import("../../src/lib/usage/completedRequestDetails.ts"); // Captured stream chunks carry a per-chunk arrival-time prefix ("[HH:MM:SS.mmm] ") // added by the request logger for streaming-latency observability (#5834). Strip it @@ -610,6 +611,120 @@ test("completedDetails cache evicts oldest entries when bounded", () => { assert.equal(usageHistory.getCompletedDetails().has(ids[ids.length - 1]), true); }); +test("JON-562 completedDetails detaches previews from large backing strings", () => { + usageHistory.clearPendingRequests(); + + const largeBacking = "JON-562-large-backing-" + "x".repeat(4 * 1024 * 1024); + const clientRequest = { + messages: [{ role: "user", content: `${largeBacking.slice(0, 1200)}...` }], + }; + const detail = { + id: "JON-562-detached", + model: "gpt-4.1", + provider: "openai", + connectionId: "conn-JON-562-detached", + startedAt: Date.now(), + clientRequest, + }; + + assert.equal(completedRequestDetails.storeCompletedDetail(detail), true); + const stored = usageHistory.getCompletedDetails().get(detail.id); + + assert.ok(stored); + assert.notStrictEqual(stored, detail, "the cache must own a detached detail object"); + assert.notStrictEqual( + stored.clientRequest, + clientRequest, + "nested payload objects must not retain the pending-request object graph" + ); +}); + +test("JON-562 cache budget wording does not claim a hard process memory ceiling", () => { + const source = fs.readFileSync( + new URL("../../src/lib/usage/completedRequestDetails.ts", import.meta.url), + "utf8" + ); + assert.match(source, /estimated cache payload budget/i); + assert.doesNotMatch(source, /hard process-wide ceiling/i); +}); + +test("JON-562 completedDetails enforces a byte budget as well as the 256-entry cap", () => { + usageHistory.clearPendingRequests(); + const chunkBytes = Math.floor(completedRequestDetails.MAX_COMPLETED_DETAILS_BYTES * 0.6); + + for (let index = 0; index < 2; index += 1) { + assert.equal( + completedRequestDetails.storeCompletedDetail({ + id: `JON-562-budget-${index}`, + model: "gpt-4.1", + provider: "openai", + connectionId: "conn-JON-562-budget", + startedAt: Date.now() + index, + streamChunks: { client: ["x".repeat(chunkBytes)] }, + }), + true + ); + } + + assert.ok( + completedRequestDetails.getCompletedDetailsByteSize() <= + completedRequestDetails.MAX_COMPLETED_DETAILS_BYTES + ); + assert.equal(usageHistory.getCompletedDetails().has("JON-562-budget-0"), false); + assert.equal(usageHistory.getCompletedDetails().has("JON-562-budget-1"), true); +}); + +test("JON-562 oversized completed detail is rejected and replacement clears its timer", () => { + usageHistory.clearPendingRequests(); + const id = "JON-562-oversized-replacement"; + const base = { + id, + model: "gpt-4.1", + provider: "openai", + connectionId: "conn-JON-562-oversized-replacement", + startedAt: Date.now(), + }; + + assert.equal(completedRequestDetails.storeCompletedDetail(base), true); + completedRequestDetails.scheduleCompletedDetailCleanup(id); + assert.deepEqual(completedRequestDetails.getCompletedDetailsCacheStats(), { + entries: 1, + cleanupTimers: 1, + bytes: completedRequestDetails.getCompletedDetailsByteSize(), + }); + + const oversized = "x".repeat(completedRequestDetails.MAX_COMPLETED_DETAILS_BYTES + 1); + assert.equal( + completedRequestDetails.storeCompletedDetail({ + ...base, + streamChunks: { client: [oversized] }, + }), + false + ); + assert.equal(usageHistory.getCompletedDetails().has(id), false); + assert.deepEqual(completedRequestDetails.getCompletedDetailsCacheStats(), { + entries: 0, + cleanupTimers: 0, + bytes: 0, + }); +}); + +test("JON-562 finalize schedules no cleanup timer when an oversized detail is not stored", () => { + usageHistory.clearPendingRequests(); + const model = "gpt-4.1"; + const provider = "openai"; + const connectionId = "conn-JON-562-oversized-finalize"; + const id = usageHistory.trackPendingRequest(model, provider, connectionId, true); + assert.ok(id); + usageHistory.updatePendingRequestStreamChunks(model, provider, connectionId, { + client: ["x".repeat(completedRequestDetails.MAX_COMPLETED_DETAILS_BYTES + 1)], + }); + + assert.equal(usageHistory.finalizePendingRequestById(id, { status: 200 }), true); + assert.equal(usageHistory.getCompletedDetails().has(id), false); + assert.equal(completedRequestDetails.getCompletedDetailsCacheStats().cleanupTimers, 0); +}); + test("streamChunks in completedDetails survives beyond the logs polling window", async () => { usageHistory.clearPendingRequests(); diff --git a/tests/unit/build/backend-only-smoke-workflows.test.ts b/tests/unit/build/backend-only-smoke-workflows.test.ts index d6de5173574..f5946cb8e63 100644 --- a/tests/unit/build/backend-only-smoke-workflows.test.ts +++ b/tests/unit/build/backend-only-smoke-workflows.test.ts @@ -9,6 +9,7 @@ // step legitimately ships the full dashboard UI in the published npm package. import test from "node:test"; import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; import fs from "node:fs"; import path from "node:path"; import * as yaml from "js-yaml"; @@ -32,6 +33,12 @@ interface WorkflowDoc { } const WORKFLOWS_DIR = path.join(process.cwd(), ".github", "workflows"); +const RELEASE_PROFILE_GUARD = path.join( + process.cwd(), + "scripts", + "build", + "assert-full-release-profile.mjs" +); function loadWorkflow(fileName: string): WorkflowDoc { const raw = fs.readFileSync(path.join(WORKFLOWS_DIR, fileName), "utf8"); @@ -51,6 +58,49 @@ test("full build remains the default when no backend-only profile is set", () => assert.equal(isBackendOnlyBuild({}), false); }); +function runReleaseProfileGuard(overrides: Record = {}) { + const env = { ...process.env }; + delete env.OMNIROUTE_BUILD_BACKEND_ONLY; + delete env.OMNIROUTE_BUILD_PROFILE; + Object.assign(env, overrides); + return spawnSync(process.execPath, [RELEASE_PROFILE_GUARD], { + cwd: process.cwd(), + env, + encoding: "utf8", + }); +} + +test("build:release checks the full-dashboard profile before deleting prior artifacts", () => { + const pkg = JSON.parse(fs.readFileSync(path.join(process.cwd(), "package.json"), "utf8")) as { + scripts?: Record; + }; + const releaseScript = pkg.scripts?.["build:release"] || ""; + assert.match( + releaseScript, + /^node scripts\/build\/assert-full-release-profile\.mjs && rm -rf \.build dist &&/, + "build:release must reject backend-only profiles before deleting the prior build" + ); +}); + +for (const env of [ + { OMNIROUTE_BUILD_BACKEND_ONLY: "1" }, + { OMNIROUTE_BUILD_PROFILE: "backend" }, + { OMNIROUTE_BUILD_PROFILE: "contributor" }, +]) { + test(`full release rejects the dashboard-stubbing profile ${JSON.stringify(env)}`, () => { + const result = runReleaseProfileGuard(env); + assert.equal(result.status, 1); + assert.match(result.stderr, /Refusing full release: .*stub the dashboard/); + }); +} + +for (const env of [{}, { OMNIROUTE_BUILD_PROFILE: "minimal" }, { UNRELATED_SETTING: "1" }]) { + test(`full release accepts a non-stubbing profile ${JSON.stringify(env)}`, () => { + const result = runReleaseProfileGuard(env); + assert.equal(result.status, 0, result.stderr); + }); +} + // jobName: null selector means "any job" — used when a file has exactly one // "Build CLI bundle" step but we don't want to hardcode/duplicate the job key. interface Target { @@ -87,7 +137,10 @@ test("npm-publish.yml 'Build CLI bundle (standalone app)' step must NOT be backe const publishJob = Object.values(doc.jobs).find((job) => job.steps.some((s) => s.name === "Build CLI bundle (standalone app)") ); - assert.ok(publishJob, "npm-publish.yml must have a job with a 'Build CLI bundle (standalone app)' step"); + assert.ok( + publishJob, + "npm-publish.yml must have a job with a 'Build CLI bundle (standalone app)' step" + ); const step = publishJob!.steps.find((s) => s.name === "Build CLI bundle (standalone app)")!; assert.equal( isBackendOnly(step), diff --git a/tests/unit/call-log-artifact-worker.test.ts b/tests/unit/call-log-artifact-worker.test.ts index 84cd2c241a2..d1db4982dfd 100644 --- a/tests/unit/call-log-artifact-worker.test.ts +++ b/tests/unit/call-log-artifact-worker.test.ts @@ -7,8 +7,13 @@ import path from "node:path"; const TEST_DATA_DIR = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-call-log-worker-")); process.env.DATA_DIR = TEST_DATA_DIR; -const { writeCallArtifactAsync, closeCallLogArtifactWriter, resolveCallLogArtifactWorker } = - await import("../../src/lib/usage/callLogArtifactWriter.ts"); +const { + writeCallArtifactAsync, + closeCallLogArtifactWriter, + getCallLogArtifactQueueStats, + resetCallLogArtifactWriterForTest, + resolveCallLogArtifactWorker, +} = await import("../../src/lib/usage/callLogArtifactWriter.ts"); test.after(async () => { await closeCallLogArtifactWriter(); @@ -53,6 +58,32 @@ function buildArtifact(id: string) { }; } +test("artifact queue fails open before retaining payloads above its byte budget", async () => { + const originalLimit = process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES; + process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES = "2048"; + + try { + const oversized = { + ...buildArtifact("queue-byte-overflow"), + requestBody: { content: "x".repeat(4096) }, + }; + + assert.equal(await writeCallArtifactAsync(oversized), null); + assert.deepEqual(getCallLogArtifactQueueStats(), { + queuedJobs: 0, + activeJobs: 0, + queuedBytes: 0, + maxQueuedJobs: 128, + maxQueuedBytes: 2048, + closing: false, + hasActiveWorker: false, + }); + } finally { + if (originalLimit === undefined) delete process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES; + else process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES = originalLimit; + } +}); + test("worker resolution covers npm, standalone, source, and missing layouts", () => { const layoutRoot = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-worker-layout-")); const createWorker = (workerFile: string) => { @@ -147,6 +178,9 @@ test("async worker writes call-log artifact and returns matching metadata", asyn }); test("bounded queue fails open and rate-limits saturation warnings", async () => { + const originalLimit = process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES; + delete process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES; + resetCallLogArtifactWriterForTest(); const originalWarn = console.warn; let warningCount = 0; console.warn = () => { @@ -164,5 +198,7 @@ test("bounded queue fails open and rate-limits saturation warnings", async () => assert.ok(results.every((result) => result === null)); } finally { console.warn = originalWarn; + if (originalLimit === undefined) delete process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES; + else process.env.CALL_LOG_ARTIFACT_MAX_QUEUED_BYTES = originalLimit; } }); diff --git a/tests/unit/call-log-cap.test.ts b/tests/unit/call-log-cap.test.ts index baa68e57b82..4bfc04d1763 100644 --- a/tests/unit/call-log-cap.test.ts +++ b/tests/unit/call-log-cap.test.ts @@ -502,6 +502,23 @@ test("getCallLogById marks missing artifacts explicitly and clears stale DB poin assert.equal((row as CallLogRow).detail_state, "missing"); }); +test("saveCallLog omits oversized bodies before payload protection", async () => { + await callLogs.saveCallLog({ + id: "pre-protection-size-bound", + timestamp: "2026-03-31T09:04:00.000Z", + method: "POST", + path: "/v1/chat/completions", + status: 200, + model: "openai/gpt-4.1", + provider: "openai", + requestBody: { payload: "x".repeat(2 * 1024 * 1024) }, + }); + + const detail = await callLogs.getCallLogById("pre-protection-size-bound"); + assert.equal(detail?.requestBody, "[omitted: call log artifact size limit exceeded]"); + assert.equal(detail?.hasRequestBody, true); +}); + test("saveCallLog keeps large payloads out of SQLite while preserving explicit detail export", async () => { const requestBody = { payload: "x".repeat(320 * 1024) }; diff --git a/tests/unit/client-bundle-no-server-only-10692.test.ts b/tests/unit/client-bundle-no-server-only-10692.test.ts index e29ae9d2faf..a20405ed562 100644 --- a/tests/unit/client-bundle-no-server-only-10692.test.ts +++ b/tests/unit/client-bundle-no-server-only-10692.test.ts @@ -31,6 +31,7 @@ const REPO_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), ".. /** Modules that pull in Node builtins (fs/net/tls) and must never be statically reachable. */ const SERVER_ONLY = new Set([ + "open-sse/services/model.ts", "src/lib/db/core.ts", "src/lib/db/adapters/driverFactory.ts", "src/lib/db/adapters/sqljsAdapter.ts", @@ -55,7 +56,9 @@ const SKIP_DIRS = new Set(["node_modules", ".git", ".build", "dist", ".next", ". /** Resolve an import specifier to a repo-relative file, or null when it leaves the repo. */ function resolveSpecifier(fromFile: string, specifier: string): string | null { let base: string; - if (specifier.startsWith(".")) { + if (specifier.startsWith("node:")) { + return specifier; + } else if (specifier.startsWith(".")) { base = path.resolve(path.dirname(path.join(REPO_ROOT, fromFile)), specifier); } else if (specifier.startsWith("@omniroute/open-sse")) { const rest = specifier.slice("@omniroute/open-sse".length).replace(/^\//, ""); @@ -140,7 +143,9 @@ function findServerOnlyPath(entry: string): string[] | null { const trail = queue.shift()!; for (const resolved of edgesOf(trail[trail.length - 1])) { if (seen.has(resolved)) continue; - if (SERVER_ONLY.has(resolved)) return [...trail, resolved]; + if (resolved.startsWith("node:") || SERVER_ONLY.has(resolved)) { + return [...trail, resolved]; + } seen.add(resolved); queue.push([...trail, resolved]); } @@ -163,7 +168,9 @@ function walk(dir: string, acc: string[] = []): string[] { function clientEntryPoints(): string[] { return walk(path.join(REPO_ROOT, "src")).filter((file) => - /^\s*["']use client["']/m.test(fs.readFileSync(path.join(REPO_ROOT, file), "utf8").slice(0, 200)) + /^\s*["']use client["']/m.test( + fs.readFileSync(path.join(REPO_ROOT, file), "utf8").slice(0, 200) + ) ); } diff --git a/tests/unit/estimateSizeFast.test.ts b/tests/unit/estimateSizeFast.test.ts index 84893097a66..3acdf251084 100644 --- a/tests/unit/estimateSizeFast.test.ts +++ b/tests/unit/estimateSizeFast.test.ts @@ -3,6 +3,7 @@ import assert from "node:assert/strict"; const { estimateSizeFast, + estimateSizeFastResult, isSmallEnoughForSemanticCache, ESTIMATE_SIZE_BYTE_LIMIT, ESTIMATE_SIZE_NODE_BUDGET, @@ -98,7 +99,25 @@ test("estimateSizeFast respects a caller-supplied byteLimit above the 256KB defa trueTotal, "must report the true accumulated size instead of early-exiting at the default 256KB" ); - assert.ok(withCustomLimit <= oneMiB, "payload must be recognized as under the caller's own limit"); + assert.ok( + withCustomLimit <= oneMiB, + "payload must be recognized as under the caller's own limit" + ); +}); + +test("estimateSizeFastResult distinguishes node exhaustion from byte overflow", () => { + const highNodePayload = Array.from({ length: 20_000 }, () => "x"); + assert.deepEqual(estimateSizeFastResult(highNodePayload, 64 * 1024), { + status: "node-budget", + bytes: 16_383, + }); + assert.equal(estimateSizeFastResult("x".repeat(70_000), 64 * 1024).status, "byte-limit"); +}); + +test("estimateSizeFast accepts an explicit node budget for valid high-node payloads", () => { + const payload = Array.from({ length: 20_000 }, () => "x"); + assert.ok(estimateSizeFast(payload, 64 * 1024) > 64 * 1024); + assert.equal(estimateSizeFast(payload, 64 * 1024, 50_000), 20_000); }); test("estimateSizeFast node-budget fail-closed return respects a caller-supplied byteLimit", () => { diff --git a/tests/unit/messages-route-memory-profile.test.ts b/tests/unit/messages-route-memory-profile.test.ts new file mode 100644 index 00000000000..6a81344bf46 --- /dev/null +++ b/tests/unit/messages-route-memory-profile.test.ts @@ -0,0 +1,335 @@ +import assert from "node:assert/strict"; +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { spawnSync } from "node:child_process"; +import test from "node:test"; + +import { + assertNoLargeBacking, + BYTES_PER_TOKEN_EQUIVALENT, + buildClaudeContextPayload, + buildWorkerEnv, + classifyContextBackingCandidates, + cleanupWorkerSnapshot, + summarizeSettledGrowth, + withRawSnapshotCleanup, +} from "../../scripts/perf/messages-route-memory-profile.ts"; + +const SCRIPT = new URL("../../scripts/perf/messages-route-memory-profile.ts", import.meta.url); + +test("JON-562 corpus varies only context bytes and hits the exact token-equivalent size", () => { + const small = buildClaudeContextPayload(100_000); + const medium = buildClaudeContextPayload(300_000); + + assert.equal( + Buffer.byteLength(JSON.stringify(small.body), "utf8"), + 100_000 * BYTES_PER_TOKEN_EQUIVALENT + ); + assert.equal( + Buffer.byteLength(JSON.stringify(medium.body), "utf8"), + 300_000 * BYTES_PER_TOKEN_EQUIVALENT + ); + assert.deepEqual(buildClaudeContextPayload(100_000), small, "the corpus must be deterministic"); + assert.equal(small.body.messages.length, medium.body.messages.length); + assert.equal(small.body.stream, medium.body.stream); + assert.equal(small.body.model, medium.body.model); +}); + +test("JON-562 settled-growth summary uses only post-GC samples", () => { + const summary = summarizeSettledGrowth( + [ + { phase: "baseline", heapUsedBytes: 1_000 }, + { phase: "after_route", heapUsedBytes: 9_000 }, + { phase: "settled", heapUsedBytes: 1_100, requestIndex: 1 }, + { phase: "settled", heapUsedBytes: 1_300, requestIndex: 2 }, + { phase: "settled", heapUsedBytes: 1_600, requestIndex: 3 }, + ], + 400 + ); + + assert.equal(summary.settledGrowthBytes, 600); + assert.equal(summary.settledSlopeBytesPerRequest, 250); + assert.equal(summary.growthToWireRatio, 1.5); +}); + +test("JON-562 worker environment is an allowlist and drops credential-shaped variables", () => { + const env = buildWorkerEnv({ + PATH: "/usr/bin:/bin", + HOME: "/tmp/profile-home", + LANG: "C.UTF-8", + LINEAR_API_KEY: "must-not-cross", + OPENAI_API_KEY: "must-not-cross", + OMNIROUTE_MANAGEMENT_TOKEN: "must-not-cross", + }); + + assert.deepEqual(env, { + PATH: "/usr/bin:/bin", + HOME: "/tmp/profile-home", + LANG: "C.UTF-8", + NODE_ENV: "test", + APP_LOG_LEVEL: "error", + }); +}); + +test("JON-562 heap analyzer uses V8 self_size, not truncated string preview length", () => { + const marker = "JON-562-600000-CONTEXT"; + const truncatedPreview = (marker + "x".repeat(2_000)).slice(0, 1_024); + const detachedSnapshot = { + snapshot: { + meta: { + node_fields: ["type", "name", "id", "self_size", "edge_count"], + node_types: [ + ["synthetic", "string", "concatenated string"], + "string", + "number", + "number", + "number", + ], + edge_fields: ["type", "name_or_index", "to_node"], + edge_types: [["internal"], "string_or_number", "node"], + }, + }, + nodes: [1, 0, 1, 1_224, 0], + edges: [], + strings: [truncatedPreview, "first", "second", "xxxxxxxx"], + }; + const slicedBackingSnapshot = { + ...detachedSnapshot, + // Marker leaf (1 KiB) + large sibling backing (4 MiB) + concatenated-string root. + nodes: [1, 0, 1, 1_024, 0, 1, 3, 2, 4 * 1024 * 1024, 0, 2, 3, 3, 24, 2], + edges: [0, 1, 0, 0, 2, 5], + }; + + const retained = classifyContextBackingCandidates(slicedBackingSnapshot, marker, 2_000_000); + const detached = classifyContextBackingCandidates(detachedSnapshot, marker, 2_000_000); + + assert.equal(retained.largeBackingRetained, true); + assert.ok(retained.maxBackingRetainedSizeBytes >= 4 * 1024 * 1024); + assert.equal(detached.largeBackingRetained, false); + assert.equal(detached.maxBackingRetainedSizeBytes, 1_224); +}); + +test("JON-562 retention gate rejects any context-sized backing result", () => { + assert.throws( + () => assertNoLargeBacking({ largeBackingRetained: true }), + /context-sized backing string/i + ); + assert.doesNotThrow(() => assertNoLargeBacking({ largeBackingRetained: false })); +}); + +test("JON-562 raw snapshot cleanup runs when a case fails or times out", async () => { + const outputDir = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-JON-562-cleanup-")); + const snapshot = path.join(outputDir, "post-gc.heapsnapshot"); + fs.writeFileSync(snapshot, "synthetic raw snapshot"); + try { + await assert.rejects( + withRawSnapshotCleanup(snapshot, async () => { + throw new Error("simulated analyzer timeout"); + }), + /simulated analyzer timeout/ + ); + assert.equal(fs.existsSync(snapshot), false); + } finally { + fs.rmSync(outputDir, { recursive: true, force: true }); + } +}); + +test("JON-562 worker deletes raw snapshots unless a successful driver handoff owns them", () => { + const outputDir = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-JON-562-worker-cleanup-")); + const failedSnapshot = path.join(outputDir, "failed.heapsnapshot"); + const handedOffSnapshot = path.join(outputDir, "handed-off.heapsnapshot"); + fs.writeFileSync(failedSnapshot, "failed worker raw snapshot"); + fs.writeFileSync(handedOffSnapshot, "successful handoff raw snapshot"); + try { + cleanupWorkerSnapshot(failedSnapshot, { + workerComplete: false, + snapshotHandoff: true, + }); + cleanupWorkerSnapshot(handedOffSnapshot, { + workerComplete: true, + snapshotHandoff: true, + }); + assert.equal(fs.existsSync(failedSnapshot), false); + assert.equal(fs.existsSync(handedOffSnapshot), true); + } finally { + fs.rmSync(outputDir, { recursive: true, force: true }); + } +}); + +test("JON-562 standalone analyzer deletes an invalid raw snapshot on failure", () => { + const outputDir = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-JON-562-analyzer-fail-")); + const snapshot = path.join(outputDir, "post-gc.heapsnapshot"); + fs.writeFileSync(snapshot, "not valid JSON"); + try { + const result = spawnSync( + process.execPath, + [ + "--import", + "tsx/esm", + SCRIPT.pathname, + "--analyze-snapshot", + "--snapshot", + snapshot, + "--output-dir", + outputDir, + "--marker", + "JON-562-invalid", + "--minimum-self-size-bytes", + "1024", + ], + { + cwd: new URL("../..", import.meta.url), + encoding: "utf8", + env: buildWorkerEnv(process.env), + timeout: 10_000, + } + ); + assert.notEqual(result.status, 0); + assert.equal(fs.existsSync(snapshot), false); + } finally { + fs.rmSync(outputDir, { recursive: true, force: true }); + } +}); + +test("JON-562 standalone analyzer deletes raw input when required arguments are missing", () => { + const outputDir = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-JON-562-analyzer-args-")); + const snapshot = path.join(outputDir, "post-gc.heapsnapshot"); + fs.writeFileSync(snapshot, "synthetic raw snapshot"); + try { + const result = spawnSync( + process.execPath, + [ + "--import", + "tsx/esm", + SCRIPT.pathname, + "--analyze-snapshot", + "--snapshot", + snapshot, + "--output-dir", + outputDir, + ], + { + cwd: new URL("../..", import.meta.url), + encoding: "utf8", + env: buildWorkerEnv(process.env), + timeout: 10_000, + } + ); + assert.notEqual(result.status, 0); + assert.equal(fs.existsSync(snapshot), false); + } finally { + fs.rmSync(outputDir, { recursive: true, force: true }); + } +}); + +test("JON-562 standalone analyzer deletes raw input when size validation fails", () => { + const outputDir = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-JON-562-analyzer-size-")); + const snapshot = path.join(outputDir, "post-gc.heapsnapshot"); + fs.writeFileSync(snapshot, "synthetic raw snapshot"); + try { + const result = spawnSync( + process.execPath, + [ + "--import", + "tsx/esm", + SCRIPT.pathname, + "--analyze-snapshot", + "--snapshot", + snapshot, + "--output-dir", + outputDir, + "--marker", + "JON-562-invalid-size", + "--minimum-self-size-bytes", + "not-a-number", + ], + { + cwd: new URL("../..", import.meta.url), + encoding: "utf8", + env: buildWorkerEnv(process.env), + timeout: 10_000, + } + ); + assert.notEqual(result.status, 0); + assert.equal(fs.existsSync(snapshot), false); + } finally { + fs.rmSync(outputDir, { recursive: true, force: true }); + } +}); + +test( + "JON-562 real /v1/messages path rejects context-sized post-GC retention", + { timeout: 90_000 }, + () => { + const outputDir = fs.mkdtempSync(path.join(os.tmpdir(), "omniroute-JON-562-canary-")); + try { + const result = spawnSync( + process.execPath, + [ + "--expose-gc", + "--import", + "tsx/esm", + SCRIPT.pathname, + "--tokens", + "100000", + "--iterations", + "2", + "--output-dir", + outputDir, + ], + { + cwd: new URL("../..", import.meta.url), + encoding: "utf8", + env: buildWorkerEnv(process.env), + timeout: 85_000, + } + ); + + assert.equal(result.status, 0, `driver failed:\n${result.stdout}\n${result.stderr}`); + const manifest = JSON.parse( + fs.readFileSync(path.join(outputDir, "workload-manifest.json"), "utf8") + ); + const caseDir = path.join(outputDir, "context-100000"); + const rows = fs + .readFileSync(path.join(caseDir, "memory.jsonl"), "utf8") + .trim() + .split("\n") + .map((line) => JSON.parse(line)); + + assert.equal(manifest.ticket, "JON-562"); + const profileCase = manifest.cases[0]; + assert.equal(profileCase.route, "/v1/messages"); + assert.equal( + profileCase.providerCalls, + 4, + "one warmup, one allocation probe and two unprofiled measured calls" + ); + assert.equal(profileCase.allocationSampling.requestCount, 1); + assert.equal(profileCase.retentionSeries.allocationSamplingEnabled, false); + assert.equal(profileCase.conditions.concurrency, 1); + assert.equal(profileCase.conditions.cancellation, "none"); + assert.equal(profileCase.conditions.stream, true); + assert.ok(rows.some((row) => row.phase === "after_route")); + assert.ok(rows.some((row) => row.phase === "settled")); + for (const row of rows.filter((entry) => entry.phase === "settled")) { + assert.equal(row.admission.activeHeavy, 0); + assert.equal(row.admission.activeHealthyHeadroom, 0); + assert.equal(row.admission.inflightBytes, 0); + assert.equal(row.admission.waiting, 0); + } + const retaining = JSON.parse( + fs.readFileSync(path.join(caseDir, "retainers.redacted.json"), "utf8") + ); + assert.equal( + retaining.largeBackingRetained, + false, + "the real production route must fail if a context-sized backing string survives forced GC" + ); + assert.equal(fs.existsSync(path.join(caseDir, "post-gc.heapsnapshot")), false); + assert.ok(fs.statSync(path.join(caseDir, "allocation.heapprofile")).size > 0); + } finally { + fs.rmSync(outputDir, { recursive: true, force: true }); + } + } +); diff --git a/tests/unit/pack-artifact-policy.test.ts b/tests/unit/pack-artifact-policy.test.ts index d6c26b6cfb0..81686b943c2 100644 --- a/tests/unit/pack-artifact-policy.test.ts +++ b/tests/unit/pack-artifact-policy.test.ts @@ -1,6 +1,6 @@ import test from "node:test"; import assert from "node:assert/strict"; -import { readFileSync } from "node:fs"; +import { readFileSync, readdirSync } from "node:fs"; import { APP_STAGING_ALLOWED_EXACT_PATHS, @@ -107,6 +107,21 @@ test("findUnexpectedArtifactPaths flags app pack files outside the allowlist", ( assert.deepEqual(unexpectedPaths, ["dist/scripts/build/prepublish.mjs", "docs/extra.md"]); }); +test("every @omniroute subpackage is covered by the package allowlist", () => { + const subpackages = readdirSync(new URL("../../@omniroute/", import.meta.url), { + withFileTypes: true, + }) + .filter((entry) => entry.isDirectory()) + .map((entry) => `@omniroute/${entry.name}/`); + + for (const prefix of subpackages) { + assert.ok( + PACK_ARTIFACT_ALLOWED_PATH_PREFIXES.includes(prefix), + `${prefix} is packed by package.json but missing from the artifact allowlist` + ); + } +}); + test("findUnexpectedArtifactPaths flags node_modules even inside an allowed prefix", () => { // Regression guard: the allowlist grants the whole `@omniroute/opencode-provider/` // prefix, which used to authorize a nested node_modules inside it — 79 MB of @@ -268,6 +283,7 @@ test("config/i18n.json ships in the tarball: allowed, required, and in package.j test("findMissingArtifactPaths flags missing root runtime files in the tarball", () => { const missingPaths = findMissingArtifactPaths( [ + "dist/BUILD_PROFILE", "dist/server.js", "bin/omniroute.mjs", "package.json", diff --git a/tests/unit/pack-dashboard-artifact.test.ts b/tests/unit/pack-dashboard-artifact.test.ts new file mode 100644 index 00000000000..834d04c5dd4 --- /dev/null +++ b/tests/unit/pack-dashboard-artifact.test.ts @@ -0,0 +1,35 @@ +import test from "node:test"; +import assert from "node:assert/strict"; + +import { findDashboardArtifactProblems } from "../../scripts/build/pack-artifact-policy.ts"; + +const ROUTES = [ + "dist/.build/next/server/app/login/page.js", + "dist/.build/next/server/app/(dashboard)/home/page.js", + "dist/.build/next/server/app/(dashboard)/dashboard/logs/page.js", + "dist/.build/next/server/app/(dashboard)/dashboard/conversations/page.js", +]; +const MANIFESTS = ROUTES.map((routePath) => + routePath.replace(/page\.js$/, "page_client-reference-manifest.js") +); +const STATIC_ASSET = "dist/.build/next/static/chunks/2bb41rtx9w33n.js"; +const FULL_ARTIFACT = [...ROUTES, ...MANIFESTS, STATIC_ASSET]; + +test("full dashboard artifact passes the release seam", () => { + assert.deepEqual(findDashboardArtifactProblems(FULL_ARTIFACT, "full\n"), []); +}); + +test("backend-only build profile fails the release seam", () => { + assert.deepEqual(findDashboardArtifactProblems(FULL_ARTIFACT, "backend\n"), [ + 'dist/BUILD_PROFILE must be full, got "backend"', + ]); +}); + +test("missing profile, route, manifests and client assets fail the release seam", () => { + assert.deepEqual(findDashboardArtifactProblems(ROUTES.slice(0, -1), ""), [ + "missing dist/.build/next/server/app/(dashboard)/dashboard/conversations/page.js", + ...MANIFESTS.map((manifestPath) => `missing ${manifestPath}`), + "missing dashboard JavaScript assets under dist/.build/next/static/", + 'dist/BUILD_PROFILE must be full, got ""', + ]); +}); diff --git a/tests/unit/prepublish-opencode-plugin-build-skip.test.ts b/tests/unit/prepublish-opencode-plugin-build-skip.test.ts index 87461666e46..2030672e8f7 100644 --- a/tests/unit/prepublish-opencode-plugin-build-skip.test.ts +++ b/tests/unit/prepublish-opencode-plugin-build-skip.test.ts @@ -14,7 +14,7 @@ import { test } from "node:test"; import assert from "node:assert/strict"; import { execFileSync } from "node:child_process"; -import { existsSync, rmSync } from "node:fs"; +import { existsSync, readFileSync, rmSync } from "node:fs"; import { join, dirname } from "node:path"; import { fileURLToPath } from "node:url"; import { resolveLocalBinEntry, isNativeExecutable } from "../../scripts/build/buildToolRunner.mjs"; @@ -25,6 +25,29 @@ const ROOT = join(__dirname, "..", ".."); const opencodePluginSrc = join(ROOT, "@omniroute", "opencode-plugin"); const opencodePluginDist = join(opencodePluginSrc, "dist", "index.js"); const opencodePluginCjs = join(opencodePluginSrc, "dist", "index.cjs"); +const prepublishSource = readFileSync(join(ROOT, "scripts", "build", "prepublish.ts"), "utf8"); + +test("prepublish preserves backend smoke builds and rejects stale standalone artifacts", () => { + assert.match( + prepublishSource, + /const requiredBuildProfile = isBackendOnlyBuild\(\) \? "backend" : "full"/ + ); + assert.match( + prepublishSource, + /const hasMatchingStandalone\s*=\s*[\s\S]*?existsSync\(standaloneServerJs\)[\s\S]*?existsSync\(standaloneBuildProfile\)[\s\S]*?=== requiredBuildProfile/ + ); + assert.match(prepublishSource, /if \(!hasMatchingStandalone\)/); +}); + +test("prepublish repairs partial production-only plugin installs before the release build", () => { + assert.match( + prepublishSource, + /const pluginDependenciesReady\s*=\s*[\s\S]*?node_modules[\s\S]*?typescript[\s\S]*?resolveLocalBinEntry\("tsup", "tsup", opencodePluginSrc\)/ + ); + assert.match(prepublishSource, /if \(!pluginDependenciesReady\)/); + assert.match(prepublishSource, /const installArgs = \[\s*"install",\s*"--include=dev",/); + assert.match(prepublishSource, /resolveLocalBinEntry\(packageName, binName, buildToolRoot\)/); +}); test("prepublish pluginAlreadyBuilt predicate recognizes a real ESM-only tsup build (#11787)", () => { rmSync(join(opencodePluginSrc, "dist"), { recursive: true, force: true }); @@ -34,9 +57,12 @@ test("prepublish pluginAlreadyBuilt predicate recognizes a real ESM-only tsup bu // devbox that hasn't built the plugin before) needs its own install first. Mirror // prepublish.ts's own install step instead of hardcoding a `.bin/tsup` path that // only exists on a devbox someone happened to `npm install` in already. - if (!existsSync(join(opencodePluginSrc, "node_modules"))) { + if ( + !existsSync(join(opencodePluginSrc, "node_modules", "typescript", "package.json")) || + !resolveLocalBinEntry("tsup", "tsup", opencodePluginSrc) + ) { const npmEntry = resolveBundledNpmEntry("npm-cli.js"); - const installArgs = ["install", "--no-audit", "--no-fund"]; + const installArgs = ["install", "--include=dev", "--no-audit", "--no-fund"]; if (npmEntry) { execFileSync(process.execPath, [npmEntry, ...installArgs], { cwd: opencodePluginSrc, diff --git a/tests/unit/resource-pressure-admission-recovery.test.ts b/tests/unit/resource-pressure-admission-recovery.test.ts new file mode 100644 index 00000000000..d94b023b1a1 --- /dev/null +++ b/tests/unit/resource-pressure-admission-recovery.test.ts @@ -0,0 +1,344 @@ +import assert from "node:assert/strict"; +import test from "node:test"; + +import { + createResourcePressureRuntime, + getAdmissionResourcePressureSeverity, + reloadResourcePressureRuntime, + type ResourceSignals, +} from "../../open-sse/utils/resourcePressure.ts"; +import { withChatAdmission } from "../../src/shared/middleware/withChatAdmission.ts"; + +const MiB = 1024 ** 2; + +function signals(observedAtMs: number, heapUsedMb: number): ResourceSignals { + return { + observedAtMs, + v8: { heapUsedBytes: heapUsedMb * MiB, heapLimitBytes: 1_000 * MiB }, + process: { + rssBytes: 200 * MiB, + externalBytes: 10 * MiB, + arrayBuffersBytes: MiB, + availableBytes: null, + constrainedBytes: null, + }, + cgroup: { currentBytes: null, maxBytes: null, highBytes: null, fileBytes: null, events: null }, + psi: null, + }; +} + +function messagesRequest(contentType = "application/json"): Request { + const body = JSON.stringify({ model: "test", messages: [{ role: "user", content: "hello" }] }); + return new Request("http://x/v1/messages", { + method: "POST", + headers: { "content-type": contentType, "content-length": String(body.length) }, + body, + }); +} + +test("JON-563: actual /v1/messages route refreshes stale critical pressure and recovers", async (t) => { + let now = 0; + let heapUsedMb = 950; + let sampleCalls = 0; + const runtime = reloadResourcePressureRuntime({ + nowMs: () => now, + staleAfterMs: 100, + maxStaleMs: 30_000, + immediateHeapUsedMb: () => 100, + sample: async () => { + sampleCalls += 1; + return signals(now, heapUsedMb); + }, + thresholds: { + sustainedSamplesCritical: 1, + sustainedSamplesRecovery: 1, + heapAbsoluteThresholdMb: null, + }, + selfRestart: { enabled: false }, + }); + t.after(() => { + reloadResourcePressureRuntime({ selfRestart: { enabled: false } }); + }); + + runtime.check(); + await runtime.whenRefreshSettled(); + assert.equal(runtime.getObservation().state.severity, "critical"); + assert.equal(sampleCalls, 1); + + heapUsedMb = 100; + now = 3_600_000; + const messagesRoute = await import("../../src/app/api/v1/messages/route.ts"); + const warnings: string[] = []; + const originalWarn = console.warn; + console.warn = (...args: unknown[]) => warnings.push(args.map(String).join(" ")); + + let response: Response; + try { + response = await messagesRoute.POST(messagesRequest("text/plain")); + } finally { + console.warn = originalWarn; + } + + assert.equal(response.status, 415, "415 proves the request passed structural admission"); + assert.equal(sampleCalls, 2, "the rejected front door must coalesce and await one fresh sample"); + assert.equal(runtime.getObservation().state.severity, "normal"); + assert.equal( + warnings.some((warning) => warning.includes("returning 503")), + false, + "a recovery refresh must not emit a false 503 decision log" + ); +}); + +test("JON-563: concurrent stale-critical requests coalesce one recovery sample", async (t) => { + let now = 0; + let heapUsedMb = 950; + let sampleCalls = 0; + const runtime = reloadResourcePressureRuntime({ + nowMs: () => now, + staleAfterMs: 100, + maxStaleMs: 30_000, + immediateHeapUsedMb: () => 100, + sample: async () => { + sampleCalls += 1; + await Promise.resolve(); + return signals(now, heapUsedMb); + }, + thresholds: { + sustainedSamplesCritical: 1, + sustainedSamplesRecovery: 1, + heapAbsoluteThresholdMb: null, + }, + selfRestart: { enabled: false }, + }); + t.after(() => { + reloadResourcePressureRuntime({ selfRestart: { enabled: false } }); + }); + + runtime.check(); + await runtime.whenRefreshSettled(); + heapUsedMb = 100; + now = 3_600_000; + + const admitted = withChatAdmission(async () => new Response("ok")); + const responses = await Promise.all( + Array.from({ length: 20 }, () => admitted(messagesRequest())) + ); + + assert.deepEqual( + responses.map((response) => response.status), + Array.from({ length: 20 }, () => 200) + ); + assert.equal(sampleCalls, 2, "concurrent retries must share one in-flight recovery sample"); +}); + +test("JON-563: fresh critical pressure still sheds and sampler failure is bounded", async (t) => { + let now = 0; + let sampleCalls = 0; + const runtime = reloadResourcePressureRuntime({ + nowMs: () => now, + staleAfterMs: 10, + maxStaleMs: 100, + retryAfterMs: 20, + immediateHeapUsedMb: () => 100, + sample: async () => { + sampleCalls += 1; + if (sampleCalls === 1) return signals(now, 950); + throw new Error("sampler unavailable"); + }, + thresholds: { + sustainedSamplesCritical: 1, + sustainedSamplesRecovery: 1, + heapAbsoluteThresholdMb: null, + }, + selfRestart: { enabled: false }, + }); + t.after(() => { + reloadResourcePressureRuntime({ selfRestart: { enabled: false } }); + }); + + runtime.check(); + await runtime.whenRefreshSettled(); + let handlerCalls = 0; + const admitted = withChatAdmission(async () => { + handlerCalls += 1; + return new Response("ok"); + }); + const freshCritical = await admitted(messagesRequest()); + assert.equal(freshCritical.status, 503); + assert.equal(freshCritical.headers.get("Retry-After"), "2"); + assert.equal(handlerCalls, 0); + assert.equal(sampleCalls, 1, "a fresh critical sample must shed without redundant sampling"); + + now = 11; + assert.equal( + await getAdmissionResourcePressureSeverity(), + "critical", + "a failed refresh retains a still-bounded critical observation" + ); + assert.equal(sampleCalls, 2); + + now = 100; + assert.equal(await getAdmissionResourcePressureSeverity(), "critical"); + assert.equal(sampleCalls, 3); + + now = 101; + const warnings: string[] = []; + const originalWarn = console.warn; + console.warn = (...args: unknown[]) => warnings.push(args.map(String).join(" ")); + try { + assert.equal( + await getAdmissionResourcePressureSeverity(), + "normal", + "an observation fails open only after maxStaleMs is exceeded" + ); + } finally { + console.warn = originalWarn; + } + assert.equal(sampleCalls, 3, "failure backoff prevents another sample at the expiry edge"); + assert.equal( + warnings.filter((warning) => warning.includes("cached critical observation expired")).length, + 1 + ); + assert.match(warnings.join("\n"), /sampleAgeMs=101 maxStaleMs=100/); +}); + +test("JON-563: duplicate module evaluations share one process pressure runtime", async (t) => { + const moduleUrl = new URL("../../open-sse/utils/resourcePressure.ts", import.meta.url).href; + const first = await import(`${moduleUrl}?jon563=first`); + const second = await import(`${moduleUrl}?jon563=second`); + t.after(() => { + second.reloadResourcePressureRuntime({ selfRestart: { enabled: false } }); + }); + + const runtime = first.reloadResourcePressureRuntime({ + immediateHeapUsedMb: () => 100, + sample: async () => signals(77, 950), + thresholds: { sustainedSamplesCritical: 1, heapAbsoluteThresholdMb: null }, + selfRestart: { enabled: false }, + }); + runtime.check(); + await runtime.whenRefreshSettled(); + + assert.equal(second.getResourcePressureObservation().state.severity, "critical"); + assert.equal(second.checkResourcePressureGuard()?.status, 503); +}); + +test( + "JON-563: a non-settling pressure sample cannot hang structural admission", + { timeout: 250 }, + async (t) => { + let now = 0; + let sampleCalls = 0; + const runtime = reloadResourcePressureRuntime({ + nowMs: () => now, + staleAfterMs: 10, + maxStaleMs: 100, + admissionRefreshTimeoutMs: 20, + immediateHeapUsedMb: () => 100, + sample: async () => { + sampleCalls += 1; + if (sampleCalls === 1) return signals(now, 950); + return new Promise(() => {}); + }, + thresholds: { sustainedSamplesCritical: 1, heapAbsoluteThresholdMb: null }, + selfRestart: { enabled: false }, + }); + t.after(() => { + reloadResourcePressureRuntime({ selfRestart: { enabled: false } }); + }); + + runtime.check(); + await runtime.whenRefreshSettled(); + now = 11; + const startedAt = performance.now(); + + assert.equal(await getAdmissionResourcePressureSeverity(), "critical"); + assert.ok(performance.now() - startedAt < 200, "admission refresh wait must be bounded"); + assert.equal(sampleCalls, 2); + } +); + +test( + "JON-563: dispose releases refresh waiters even while the sampler is in flight", + { timeout: 250 }, + async (t) => { + let resolveSample!: (value: ResourceSignals) => void; + const pendingSample = new Promise((resolve) => { + resolveSample = resolve; + }); + t.after(() => resolveSample(signals(1, 100))); + const runtime = createResourcePressureRuntime({ + immediateHeapUsedMb: () => 100, + sample: () => pendingSample, + selfRestart: { enabled: false }, + }); + + runtime.check(); + await new Promise((resolve) => setImmediate(resolve)); + runtime.dispose(); + + await runtime.whenRefreshSettled(); + } +); + +test("JON-563: invalid reload leaves the working process runtime intact", async (t) => { + let now = 0; + let heapUsedMb = 100; + const runtime = reloadResourcePressureRuntime({ + nowMs: () => now, + staleAfterMs: 10, + immediateHeapUsedMb: () => 100, + sample: async () => signals(now, heapUsedMb), + thresholds: { sustainedSamplesCritical: 1, heapAbsoluteThresholdMb: null }, + selfRestart: { enabled: false }, + }); + t.after(() => { + reloadResourcePressureRuntime({ selfRestart: { enabled: false } }); + }); + + runtime.check(); + await runtime.whenRefreshSettled(); + assert.throws(() => reloadResourcePressureRuntime({ staleAfterMs: -1 }), /staleAfterMs/); + + heapUsedMb = 950; + now = 11; + await getAdmissionResourcePressureSeverity(); + await runtime.whenRefreshSettled(); + assert.equal(await getAdmissionResourcePressureSeverity(), "critical"); +}); + +test("JON-563: a scheduler that never starts work cannot hang structural admission", async (t) => { + let now = 0; + let heapUsedMb = 950; + let sampleCalls = 0; + let scheduleCalls = 0; + const runtime = reloadResourcePressureRuntime({ + nowMs: () => now, + staleAfterMs: 10, + maxStaleMs: 100, + admissionRefreshTimeoutMs: 20, + immediateHeapUsedMb: () => 100, + schedule: (refresh) => { + scheduleCalls += 1; + if (scheduleCalls === 1) setImmediate(refresh); + }, + sample: async () => { + sampleCalls += 1; + return signals(now, heapUsedMb); + }, + thresholds: { sustainedSamplesCritical: 1, heapAbsoluteThresholdMb: null }, + selfRestart: { enabled: false }, + }); + t.after(() => { + reloadResourcePressureRuntime({ selfRestart: { enabled: false } }); + }); + + runtime.check(); + await runtime.whenRefreshSettled(); + + heapUsedMb = 100; + now = 11; + assert.equal(await getAdmissionResourcePressureSeverity(), "critical"); + assert.equal(sampleCalls, 1, "the dead scheduler cannot start a second sample"); + assert.equal(scheduleCalls, 2); +});