perf(inference): reuse selected chat capability - #6730
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughThe PR adds a per-run, one-shot cache for successful OpenAI Chat Completions validation, propagates it through NIM onboarding, reuses matching results during smoke verification, and invalidates cached state on validation or setup failures. ChangesOnboarding capability cache
Estimated code review effort: 3 (Moderate) | ~25 minutes Sequence Diagram(s)sequenceDiagram
participant SelectionValidation
participant OnboardInferenceCapabilityCache
participant VerifyOnboardInferenceSmoke
participant SmokeProbe
SelectionValidation->>OnboardInferenceCapabilityCache: Record successful Chat Completions validation
VerifyOnboardInferenceSmoke->>OnboardInferenceCapabilityCache: Consume matching validation
OnboardInferenceCapabilityCache-->>VerifyOnboardInferenceSmoke: Return reusable result
VerifyOnboardInferenceSmoke->>VerifyOnboardInferenceSmoke: Skip duplicate smoke probe
VerifyOnboardInferenceSmoke->>SmokeProbe: Run probe when no reusable result exists
SmokeProbe-->>OnboardInferenceCapabilityCache: Invalidate after probe failure
Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
E2E Advisor RecommendationRequired E2E: Dispatch hint: Full advisor summaryE2E Recommendation AdvisorBase: Required E2E
Optional E2E
New E2E recommendations
Dispatch hint
|
ef6a203 to
93e94c9
Compare
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/onboard/inference-capability-cache.test.ts`:
- Around line 21-27: Remove both conditional request increments in the cache
test and assert the boolean results of takeCompletedOpenAiChat directly,
preserving the expected first and second one-shot outcomes and request counts.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: f33d1909-75e5-4467-aeba-955caf4b8cfa
📒 Files selected for processing (13)
src/lib/inference/onboard-probes.tssrc/lib/onboard/inference-capability-cache.test.tssrc/lib/onboard/inference-capability-cache.tssrc/lib/onboard/inference-providers/remote.tssrc/lib/onboard/inference-providers/types.tssrc/lib/onboard/inference-selection-validation.test.tssrc/lib/onboard/inference-selection-validation.tssrc/lib/onboard/machine/handlers/provider-inference.tssrc/lib/onboard/setup-inference.tssrc/lib/onboard/setup-nim-flow.tssrc/lib/onboard/setup-nim-selection.tstest/helpers/onboard-smoke-verifier-harness.tstest/onboard-smoke-verifier.test.ts
|
✨ Thanks for the performance work, @HOYALIM. Adding a one-shot capability cache to avoid repeated Chat Completions requests during onboarding could reduce latency. Ready for maintainer review. |
93e94c9 to
231eca3
Compare
There was a problem hiding this comment.
🧹 Nitpick comments (1)
src/lib/onboard/setup-nim-flow.test.ts (1)
284-286: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winPrefer a behavioral cache assertion over
toBeInstanceOf.This couples the setup-flow test to the concrete cache class without proving that the returned cache is usable or correctly propagated. Exercise its public behavior, or cover propagation at the public onboarding boundary, while keeping this test focused on observable setup results.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/lib/onboard/setup-nim-flow.test.ts` around lines 284 - 286, Update the assertion in the setup-flow test around resultWithoutCache to verify the returned inferenceCapabilityCache through its public behavior rather than checking its concrete OnboardInferenceCapabilityCache type. Exercise a meaningful cache operation or assert propagation through the public onboarding result, while keeping the existing observable setup-result assertions unchanged.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@src/lib/onboard/setup-nim-flow.test.ts`:
- Around line 284-286: Update the assertion in the setup-flow test around
resultWithoutCache to verify the returned inferenceCapabilityCache through its
public behavior rather than checking its concrete
OnboardInferenceCapabilityCache type. Exercise a meaningful cache operation or
assert propagation through the public onboarding result, while keeping the
existing observable setup-result assertions unchanged.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 5793c17f-7c95-4fed-8fc6-0130d1475eab
📒 Files selected for processing (14)
src/lib/inference/onboard-probes.tssrc/lib/onboard/inference-capability-cache.test.tssrc/lib/onboard/inference-capability-cache.tssrc/lib/onboard/inference-providers/remote.tssrc/lib/onboard/inference-providers/types.tssrc/lib/onboard/inference-selection-validation.test.tssrc/lib/onboard/inference-selection-validation.tssrc/lib/onboard/machine/handlers/provider-inference.tssrc/lib/onboard/setup-inference.tssrc/lib/onboard/setup-nim-flow.test.tssrc/lib/onboard/setup-nim-flow.tssrc/lib/onboard/setup-nim-selection.tstest/helpers/onboard-smoke-verifier-harness.tstest/onboard-smoke-verifier.test.ts
🚧 Files skipped from review as they are similar to previous changes (13)
- src/lib/onboard/inference-providers/types.ts
- src/lib/inference/onboard-probes.ts
- src/lib/onboard/inference-providers/remote.ts
- src/lib/onboard/setup-nim-selection.ts
- src/lib/onboard/inference-capability-cache.test.ts
- test/onboard-smoke-verifier.test.ts
- src/lib/onboard/setup-inference.ts
- src/lib/onboard/machine/handlers/provider-inference.ts
- src/lib/onboard/setup-nim-flow.ts
- src/lib/onboard/inference-capability-cache.ts
- src/lib/onboard/inference-selection-validation.test.ts
- test/helpers/onboard-smoke-verifier-harness.ts
- src/lib/onboard/inference-selection-validation.ts
231eca3 to
6143acf
Compare
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/helpers/onboard-smoke-verifier-harness.ts`:
- Around line 94-108: Update the capability-cache priming in the invocation flow
around selectedChatCapability to use the effective endpoint, model, provider,
and authMode values from input rather than hard-coded values. Ensure
rememberCompletedOpenAiChat receives the same capability key that
verifyOnboardInferenceSmoke will use after applying ...input, so the selected
capability path exercises cache reuse.
- Around line 97-101: Update the cache priming call in the harness around
capabilityCache.rememberCompletedOpenAiChat to capture and assert its boolean
result. Fail the test harness immediately when the call returns false, ensuring
the probe mock’s successful response actually primes the cache before reuse is
tested.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: c406b73a-5242-478e-9ec6-b0119c4d4564
📒 Files selected for processing (14)
src/lib/inference/onboard-probes.tssrc/lib/onboard/inference-capability-cache.test.tssrc/lib/onboard/inference-capability-cache.tssrc/lib/onboard/inference-providers/remote.tssrc/lib/onboard/inference-providers/types.tssrc/lib/onboard/inference-selection-validation.test.tssrc/lib/onboard/inference-selection-validation.tssrc/lib/onboard/machine/handlers/provider-inference.tssrc/lib/onboard/setup-inference.tssrc/lib/onboard/setup-nim-flow.test.tssrc/lib/onboard/setup-nim-flow.tssrc/lib/onboard/setup-nim-selection.tstest/helpers/onboard-smoke-verifier-harness.tstest/onboard-smoke-verifier.test.ts
🚧 Files skipped from review as they are similar to previous changes (12)
- src/lib/onboard/inference-capability-cache.test.ts
- src/lib/onboard/inference-selection-validation.test.ts
- test/onboard-smoke-verifier.test.ts
- src/lib/onboard/setup-nim-flow.test.ts
- src/lib/onboard/inference-selection-validation.ts
- src/lib/onboard/inference-providers/types.ts
- src/lib/onboard/setup-nim-selection.ts
- src/lib/onboard/setup-inference.ts
- src/lib/inference/onboard-probes.ts
- src/lib/onboard/inference-providers/remote.ts
- src/lib/onboard/setup-nim-flow.ts
- src/lib/onboard/machine/handlers/provider-inference.ts
Signed-off-by: Ho Lim <subhoya@gmail.com>
6143acf to
ecd060b
Compare
<!-- markdownlint-disable MD041 --> ## Summary Release-prep documentation for v0.0.82 now summarizes user-facing changes merged since v0.0.81. It also closes stale wording in the stopped-sandbox backup, snapshot-clone, Ollama selection, and custom-policy authoring guidance. ## Changes - Add the `v0.0.82` section to `docs/about/release-notes.mdx` with links to the focused user guides. - Document that snapshot clones receive a destination-owned dashboard port before destructive replacement begins. - Align `backup-all` guidance with eligible stopped Docker-driver sandboxes that NemoClaw starts temporarily. - Describe the running and stopped Ollama menu states without claiming one fixed label. - Document runtime rejection of catch-all hosts in custom policy files. ### Source summary - [#6748](#6748) -> `docs/about/release-notes.mdx`, `docs/manage-sandboxes/lifecycle.mdx`, and `docs/reference/commands.mdx`: Summarize non-destructive sandbox `stop` and `start` commands. - [#6723](#6723) -> `docs/about/release-notes.mdx`, `docs/manage-sandboxes/backup-restore.mdx`, and `docs/reference/commands.mdx`: Record temporary startup and cleanup for eligible stopped-sandbox backups. - [#6749](#6749) -> `docs/about/release-notes.mdx` and `docs/manage-sandboxes/backup-restore.mdx`: Document destination-owned dashboard ports for snapshot clones. - [#6764](#6764) -> `docs/about/release-notes.mdx`: Summarize installer handling of route-only onboarding placeholders. - [#6771](#6771) -> `docs/about/release-notes.mdx`, `docs/inference/set-up-vllm.mdx`, `docs/inference/choose-inference-provider.mdx`, `docs/reference/commands.mdx`, and `docs/reference/platform-support.mdx`: Summarize managed-vLLM storage gates, immutable image digests, and the explicit override boundary. - [#6759](#6759) -> `docs/about/release-notes.mdx`: Record early, actionable OpenShell gateway-port conflict diagnostics. - [#6753](#6753) -> `docs/about/release-notes.mdx` and `docs/inference/set-up-ollama.mdx`: Document truthful running and stopped Ollama menu states. - [#6776](#6776) -> `docs/about/release-notes.mdx`: Summarize proxy-independent loopback readiness checks. - [#6769](#6769) -> `docs/about/release-notes.mdx`: Record compatible endpoint and agent guidance when Chat Completions is unavailable. - [#6730](#6730) -> `docs/about/release-notes.mdx`: Summarize bounded reuse of an eligible successful Chat Completions check. - [#6768](#6768) -> `docs/about/release-notes.mdx`: Record route-reservation repair during resumed onboarding. - [#6742](#6742) -> `docs/about/release-notes.mdx`: Summarize pre-mutation resolution of secret-free sandbox create intent. - [#6721](#6721) -> `docs/about/release-notes.mdx` and `docs/get-started/quickstart-langchain-deepagents-code.mdx`: Record bounded cleanup of completed managed Deep Agents headless sessions. - [#6731](#6731) -> `docs/about/release-notes.mdx` and `docs/network-policy/customize-network-policy.mdx`: Document runtime rejection of catch-all custom-policy destinations. - [#6729](#6729) -> `docs/about/release-notes.mdx` and `docs/get-started/prerequisites.mdx`: Record the Node.js 22.19 minimum. - [#6735](#6735) -> `docs/about/release-notes.mdx` and `docs/reference/platform-support.mdx`: Summarize the Ubuntu 26.04 userspace contract without claiming pending host or live validation. - [#6775](#6775) -> `docs/about/release-notes.mdx` and `docs/resources/community-contributions.mdx`: Route independent solutions outside canonical supported-product documentation. - [#6740](#6740) -> `docs/about/release-notes.mdx`: Summarize the semantic dependency-upgrade contributor workflow. - [#6777](#6777) -> `docs/about/release-notes.mdx` and `docs/CONTRIBUTING.md`: Summarize the route-safe documentation-refactor workflow. - [#6741](#6741) -> `docs/about/release-notes.mdx` and `docs/security/openclaw-2026.6.10-dependency-review.md`: Summarize reviewed npm archive verification and audit enforcement. - [#6739](#6739) -> `docs/about/release-notes.mdx` and `docs/security/openclaw-2026.6.10-dependency-review.md`: Record the locked offline dependency graph for the managed OpenClaw WeChat runtime. - [#6737](#6737) -> `docs/about/release-notes.mdx`: Record removal of the messaging build plan from final OpenClaw and Hermes image environments. - [#6733](#6733) -> `docs/about/release-notes.mdx`: Summarize cached plugin dependency layers for source and blueprint rebuilds. ### Skipped from docs-skip - None. No commit or changed path in `v0.0.81..origin/main` matched `openclaw-sandbox-permissive.yaml` or `config-show`, and the drafted content contains none of the configured skip terms. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [x] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [ ] Tests added or updated for changed behavior - [ ] Existing tests cover changed behavior — justification: - [x] Tests not applicable — justification: This is a documentation-only release-prep update; behavior is protected by the merged source PRs, and the documentation build validates the changed routes and agent variants. - [x] Docs updated for user-facing behavior changes - [ ] Docs not applicable — justification: - [ ] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [ ] Sensitive-path review completed or maintainer-approved waiver recorded — reviewer/approval link/justification: - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: ## Verification - [x] PR description includes the DCO sign-off declaration and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run check:diff` passed when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — tests are not applicable for this documentation-only change. - [ ] Applicable broad gate passed — `npm test` for broad runtime/test-harness changes; `npm run check` for repo-wide validation/coverage changes — command/result: not run for this documentation-only change. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) — 0 errors; two pre-existing Fern warnings remain. - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) — no new pages. --- Signed-off-by: Charan Jagwani <cjagwani@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated release notes with improvements to sandbox recovery, onboarding, session management, policy validation, storage checks, and system requirements. * Clarified Ollama setup instructions and status labels. * Documented safer snapshot restoration, including dedicated ports and protection against destructive failures. * Expanded `backup-all` coverage to include eligible stopped sandboxes. * Added guidance rejecting broad or catch-all network destinations in custom policies. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>
Summary
Avoid repeating the same successful Chat Completions capability request during one onboarding run. Provider selection records a narrowly scoped, one-shot capability receipt that the immediately following host smoke consumes only when endpoint, model, authentication mode, and semantic requirements still match.
Changes
Verification
npx vitest run --project cli src/lib/onboard/inference-capability-cache.test.ts src/lib/onboard/inference-selection-validation.test.tsnpx vitest run --project integration test/onboard-smoke-verifier.test.tsnpm run build:clinpm run typecheck:clinpm run check:diffSigned-off-by: Ho Lim subhoya@gmail.com
Summary by CodeRabbit