fix(inference): recompute context window on model switch - #5457
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (6)
📝 WalkthroughWalkthroughFixes stale context window propagation on ChangesContext Window Recomputation on Model Switch
Sequence Diagram(s)sequenceDiagram
participant CLI as nemoclaw inference set
participant runInferenceSet
participant resolveContextWindowForModel
participant OllamaRuntime as Ollama daemon
participant VllmEndpoint as vLLM /v1/models
CLI->>runInferenceSet: provider, model, sandbox
runInferenceSet->>resolveContextWindowForModel: provider, model
alt ollama-local
resolveContextWindowForModel->>OllamaRuntime: warm model
resolveContextWindowForModel->>OllamaRuntime: probe context length
OllamaRuntime-->>resolveContextWindowForModel: number or null
else vllm-local
resolveContextWindowForModel->>VllmEndpoint: GET /v1/models
VllmEndpoint-->>resolveContextWindowForModel: max_model_len number or null
else cloud
resolveContextWindowForModel-->>runInferenceSet: DEFAULT 131072
end
resolveContextWindowForModel-->>runInferenceSet: contextSize or null
alt resolved number
runInferenceSet->>runInferenceSet: log token count
runInferenceSet->>runInferenceSet: patchOpenClawInferenceConfig with contextSize
else null
runInferenceSet->>runInferenceSet: warn indeterminate, keep existing
runInferenceSet->>runInferenceSet: patchOpenClawInferenceConfig unchanged
end
runInferenceSet-->>CLI: updated sandbox config
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage in the Show a code coverage summary of the most covered files.
TypeScript / code-coverage/cliThe overall coverage in the Show a code coverage summary of the most covered files.
Updated |
E2E Advisor RecommendationRequired E2E: Dispatch hint: Full advisor summaryE2E Recommendation AdvisorBase: Required E2E
Optional E2E
New E2E recommendations
Dispatch hint
|
Vitest E2E Scenario RecommendationRequired Vitest E2E scenarios: Dispatch required Vitest E2E scenarios:
Full Vitest E2E advisor summaryVitest E2E Scenario AdvisorBase: Required Vitest E2E scenarios
Optional Vitest E2E scenarios
Relevant changed files
|
PR Review AdvisorFindings: 1 needs attention, 2 worth checking, 0 nice ideas Review findings🛠️ Needs attention
🔎 Worth checking
🌱 Nice ideas
Consider writing more tests for
Since last review detailsCurrent findings:
This is an automated advisory review. A human maintainer must make the final merge decision. |
jason-ma-nv
left a comment
There was a problem hiding this comment.
Reviewed this against #5458 (the auto-generated PR for the same issue, now closed) — this is the stronger fix and the one to merge.
Why this one: resolveOllamaRuntimeContextWindow only reports a window when the model is already loaded. This PR warms the model first (warmOllamaModel) and then probes, so the canonical repro from #5456 (cloud → ollama-local on a freshly-pulled model) actually detects and writes the runtime window instead of leaving it to chance. The shared resolveVllmContextWindowFromModels parser (used by both onboard and inference set) and the per-switch log/warning are also nice — the warning pointing users to rebuild directly addresses the missing-feedback note in the issue.
Three things worth grafting from #5458 before merge:
-
Floor the Ollama value at the agent minimum. onboard does
Math.max(detected, MIN_AUTODETECTED_OLLAMA_CONTEXT_WINDOW)(16384).resolveContextWindowForModelcurrently writes the raw probed value, so a sub-16k model would land below onboard's floor and diverge from the onboard path the issue asks us to mirror. -
Add timeouts to the vLLM probe.
probeVllmContextWindowrunscurl -sf http://127.0.0.1:<port>/v1/modelswith no--connect-timeout/--max-time. If the port is open but the server is slow/wedged,inference setcan hang. #5458 used--connect-timeout 3 --max-time 5with validated curl args. -
Document the new behavior. #5458 added a paragraph to
docs/inference/switch-inference-providers.mdxstating thatinference setrecomputes the window per provider on a switch. Worth pulling in (drop the stray$$prefix it had in the code span).
Minor: the fix commit has an empty body and the PR description has CRLF line endings — not blocking.
Note: I verified the warm/probe and floor semantics from the source rather than a live Ollama run, so the Ollama detection path is worth a quick manual sanity check on a host with Ollama before merge.
## Summary Refreshes release-prep documentation for NemoClaw v0.0.65. Adds the v0.0.65 release-notes section and refreshes generated `nemoclaw-user-*` skills from the Fern MDX source docs. ## Changes - Added the v0.0.65 release notes to `docs/about/release-notes.mdx` with links to the deeper docs pages for lifecycle, troubleshooting, inference, CLI commands, messaging, credentials, network policy, Hermes, and sub-agents. - Regenerated the `nemoclaw-user-*` skills with `scripts/docs-to-skills.py` so release-prep skill output matches the merged source docs. - Used the v0.0.65 announcement discussion as release context: #5472. ## Source Summary - #2492 -> `docs/about/release-notes.mdx`: Documents deadline-based gateway wait reliability in the v0.0.65 recovery summary. - #4958 -> `docs/about/release-notes.mdx`: Documents re-execed OpenClaw gateway health check recovery in the sandbox recovery summary. - #5163 -> `docs/about/release-notes.mdx`: Documents safer uninstall TTY confirmation behavior in the day-two CLI summary. - #5178 -> `docs/about/release-notes.mdx`: Documents fail-closed config restore merge behavior in the rebuild and restore summary. - #5179 -> `docs/about/release-notes.mdx`: Documents WeChat QR token redaction in the messaging summary. - #5182 -> `docs/about/release-notes.mdx`: Documents sustained gateway serving checks in the recovery summary. - #5194 -> `docs/about/release-notes.mdx`: Documents model-router teardown during uninstall in the day-two CLI summary. - #5195 -> `docs/about/release-notes.mdx`: Documents Shields auto-restore lock reconfirmation in the rebuild and restore summary. - #5198 -> `docs/about/release-notes.mdx`: Documents Docker Desktop WSL CDI injection failure handling in the onboarding diagnostics summary. - #5201 -> `docs/about/release-notes.mdx`: Documents sandbox download/upload wrappers and sessions export in the day-two CLI summary. - #5205 -> `docs/about/release-notes.mdx`: Documents reporter-owned model metadata preservation in the rebuild and restore summary. - #5214 -> `docs/about/release-notes.mdx`: Documents managed vLLM model preflight before side effects in the inference setup summary. - #5215 -> `docs/about/release-notes.mdx`: Documents managed vLLM extra serve arguments in the inference setup summary. - #5216 -> `docs/about/release-notes.mdx`: Documents silent OpenClaw runtime fallback surfacing in the onboarding diagnostics summary. - #5225 -> `docs/about/release-notes.mdx`: Documents persisted sandbox gateway lookup in the gateway recovery summary. - #5238 -> `docs/about/release-notes.mdx`: Documents sub-agent gateway dial-back through the sandbox interface in the Hermes and sub-agent summary. - #5248 -> `docs/about/release-notes.mdx`: Documents Discord per-account proxy resolution in the messaging summary. - #5264 -> `docs/about/release-notes.mdx`: Documents reserved Hermes port `8642` handling in the Hermes compatibility summary. - #5267 -> `docs/about/release-notes.mdx`: Documents the narrower Hermes baseline policy in the Hermes compatibility summary. - #5321 -> `docs/about/release-notes.mdx`: Documents restored gateway guard chains in the gateway recovery summary. - #5328 -> `docs/about/release-notes.mdx`: Documents compact persisted messaging plans in the messaging summary. - #5338 -> `docs/about/release-notes.mdx`: Documents manifest channel migration in the messaging summary. - #5352 -> `docs/about/release-notes.mdx`: Documents persisted agent preservation through registry recovery in the rebuild and restore summary. - #5371 -> `.agents/skills/nemoclaw-user-reference/references/commands.md`: Refreshes generated skill output for custom build cache and layer-ordering source docs. - #5379 -> `docs/about/release-notes.mdx`: Documents dashboard port allocation across multiple NemoClaw gateways in the recovery summary. - #5382 -> `docs/about/release-notes.mdx`: Documents recovery when an active gateway has no sandbox spec in the recovery summary. - #5389 -> `.agents/skills/nemoclaw-user-reference/references/troubleshooting.md`: Refreshes generated skill output for declared agent `forward_ports` recovery source docs. - #5400 -> `docs/about/release-notes.mdx`: Documents bounded compatible endpoint probes in the inference setup summary. - #5410 -> `docs/about/release-notes.mdx`: Documents provider credential hash removal from sandbox registry entries in the messaging summary. - #5418 -> `docs/about/release-notes.mdx`: Documents summarized inference validation failures in the onboarding diagnostics summary. - #5457 -> `docs/about/release-notes.mdx`: Documents context-window recomputation after runtime model switches in the inference setup summary. - #5463 -> `docs/about/release-notes.mdx`: Documents cleanup of hard-coded messaging channel stragglers in the messaging summary. ## Skipped - #5366 matched `docs/.docs-skip` entries through skipped experimental paths, so this PR does not add new release-note text for that commit. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [x] Doc only (includes code sample changes) ## Verification - [x] Git hooks passed during commit and push, or `npx prek run --from-ref main --to-ref HEAD` passes - [ ] Targeted tests pass for changed behavior - [ ] Full `npm test` passes (broad runtime changes only) - [ ] Tests added or updated for new or changed behavior - [x] No secrets, API keys, or credentials committed - [x] Docs updated for user-facing behavior changes - [ ] `npm run docs` builds without warnings (doc changes only) - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) Verification notes: - `npm run docs` passed after rerunning outside the sandbox. Fern reported 0 errors and 1 hidden warning. - The first sandboxed `npm run docs` attempt failed before validation because `tsx` could not create its local IPC pipe under sandbox restrictions. - `npm run build:cli` passed before push to refresh the local `dist/` artifacts used by the CLI typecheck hook. - `npm test` was not run because this is a docs-only release refresh. --- Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Released NemoClaw v0.0.65 with improved gateway/sandbox recovery, safer day-two workflows, and enhanced Hermes compatibility. * Added managed vLLM extra-arguments configuration via `NEMOCLAW_VLLM_EXTRA_ARGS_JSON`. * Added Hermes troubleshooting guidance for port forwarding and health checks. * **Documentation** * Updated NVIDIA Endpoints/NIM setup and examples to use `NVIDIA_INFERENCE_API_KEY`. * Refined NVIDIA network policy and Model Router API base configuration. * Expanded CLI/environment variable documentation (including sub-agent gateway connectivity) and plugin build performance tips. * **Tests** * Expanded Vitest-backed E2E release validation coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
Summary
nemoclaw inference setkept the previous model's context window when switching, for any provider (patchOpenClawInferenceConfigcloned the prior model entry, includingcontextWindow). This recomputes the window for the target model — Ollama warm+probe, vLLM/v1/modelsmax_model_len, cloud the onboard default — so the in-sandbox config matches the model instead of carrying a stale window (e.g. a 131072 cloud default kept for an Ollama model whose runtime is ~16k → silent overflow).Related Issue
Fixes #5456
Changes
src/lib/inference/context-window.ts—resolveContextWindowForModel(provider, model):ollama-local: warm the model, then probe its runtime context length (resolveOllamaRuntimeContextWindow).vllm-local: read the running server'smax_model_lenfrom/v1/models, via a sharedresolveVllmContextWindowFromModelsextracted fromapplyVllmRuntimeContextWindowso onboard andinference setuse the same source.DEFAULT_CONTEXT_WINDOW = 131072, new constant inconfig.ts).nullwhen the local runtime is unreachable.src/lib/inference/vllm-runtime-context.ts: extractresolveVllmContextWindowFromModels(pure parse);applyVllmRuntimeContextWindownow wraps it (behavior unchanged).src/lib/actions/inference-set.ts:runInferenceSet(OpenClaw path) computes the window and passes it topatchOpenClawInferenceConfig(new optionalcontextWindowparam →buildProviderConfigsetsmodel.contextWindowinstead of inheriting). Onnullit keeps the existing value and warns to runnemoclaw <name> rebuild(which re-runs onboard and re-probes).inference setprocess.context-window.test.ts(helper dispatch incl. vLLM + null), existingvllm-runtime-context.test.ts(shared parse core),inference-set.test.ts(caller writes the recomputed window / preserves + warns on null).Out of scope — Hermes: its config has no context-window field of any name (
patchHermesInferenceConfigandagents/hermes/configwrite onlymodel.default,base_url,provider,api_key,api_mode), so this fix has no Hermes counterpart. A per-model Hermes window would be a separate config-schema change.Type of Change
Verification
npx prek run --from-ref main --to-ref HEADpassesnpm testpasses (broad runtime changes only)npm run docsbuilds without warnings (doc changes only)Signed-off-by: Hung Le hple@nvidia.com
Summary by CodeRabbit
New Features
Tests