test(e2e): add model router routed guard for #3255 - #3594
Conversation
📝 Walkthrough<review_stack_artifact_start /> Introduces WeChat host-QR login, seeding, diagnostics, policy presets, onboarding wiring, sandbox state preservation, build-image changes, new E2E scripts/tests, and workflow updates.New/updated GitHub Actions workflows (nightly/regression/scenario), selective dispatch allowlists, artifact upload changes, and scenario runner refactor for WSL/matrix and artifact handling.range_10d4ab2b44c8 range_2324a70c63db range_b4034b60d0ce range_7338775154ff range_611d8f597178 range_d2b284f6399b range_66d7b4f7e394 range_0ab00fb20706 range_f9d51c206481 range_55780eb8d95f range_e70245ca6506 range_34f4b2542f63 range_e276ad8646c6 range_c8cf35e96b4c range_5ee39dfa3923 range_3657f50bd3b8 range_71dd9df410fc range_5ee39dfa3923 range_3657f50bd3b8 range_71dd9df410fc range_0b957b5a6aa9 range_0b957b5a6aa9 range_0b957b5a6aa9 range_0b957b5a6aa9 range_3657f50bd3b8 range_71dd9df410fc range_5ee39dfa3923 range_0b957b5a6aa9Adds WeChat build ARG/ENV, copies seed script into image, installs openclaw-weixin during build and runs seeding step; Dockerfile.base cleanup; dockerBuild quiet/BuildKit handling.range_07a14ed6a3e5 range_39f97e517d78 range_9fc31e678434 range_5cabaeab45ed range_a3f4fc6ba610 range_a19812e9bee9 range_f5670a323e1d range_07f12eacd3d3 range_85b78de4574e range_9104af59967f range_635425432d55 range_f828ffdb28b6 range_782ff12085f1 range_20bc9dbecbd7 range_04ea63a1f8ab range_75918e781944Adds scripts/seed-wechat-accounts.py, documents NEMOCLAW_WECHAT_CONFIG_B64, ensures openclaw-weixin plugin entry enabled while deferring accounts seeding to seed script.range_79caf6d7406a range_1f701ebc33f1 range_583415b991c6 range_5b5de73a475f range_7d370211d0a2 range_29281d96eefb range_a9b34adef5e7 range_98bd292b758c range_4230bd603e03 range_c926be92176a range_b1a2d516965f range_07f12eacd3d3Adds runtime diagnostics script, QR bootstrap/poll client, host-side QR login orchestration, related tests, and small dependency addition for QR rendering.range_b8858e1455af range_96e0f5848e23 range_56ed0dc2c2fe range_9c3965507cc2 range_689052e1d478 range_9cf8d5f1cf78 range_faa07f3cb690 range_1f58c980a902 range_fc3a8eddf4a9 range_c191310ee617 range_b67663932ec6 range_a72bf50a1782 range_12cb61367505 range_97ffe52a4377 range_87127b1070c0 range_f5df9bce4ae4 range_c135bfb8634a range_6577388d196f range_042a71747c2b range_2dfccd4ec704 range_784421733045 range_b6f936e102d8 range_d41a9a7b6c04 range_f3df7b0e2c57 range_1c48cd66580b range_5be11fbe594e range_ec196ce7ceaf range_7ff6b845fbe5 range_677842fae91c range_4e56ccda5e96 range_b20e36d77440 range_54c113246975Wires WeChat into onboarding flow: wechat-config helpers, host-QR handler registry and dispatch, messaging-channel setup, messaging reuse/backfill, session fields (wechatConfig, disabledChannels), Dockerfile patching, and rebuild preservation.range_579a661143c7 range_263e5b4862a0 range_6a7afd0ba4d8 range_edf7f70967b6 range_069c1b86cb44 range_55c020e3b95a range_bc6e9eb34d75 range_4033344d6cac range_18769d9672aa range_9378aabb0e57 range_cb91f0a9d32e range_01f7b8ad93b8 range_b0c9677bed33 range_e7bf17c149ec range_9950a6b78b88 range_869849c6b59d range_70141a44f10f range_16ebc7c19330 range_848d7636890dStages seed script into build context, improves build failure diagnostics, adds AUDIT_SYMLINK_WHITELIST, updates pre-backup audit output parsing, and adds/exercises security tests for tar extraction and backup auditing.range_20bc9dbecbd7 range_04ea63a1f8ab range_75918e781944 range_bc7c4f1902e4 range_d79d4a7e5d8e range_86a182da0364 range_a2947ca272d2 range_c53d7cc6dc0e range_0de1591c3b86 range_53b79d943f8f range_b5c5f4aa9bff range_585c36d2c98f range_0d97583f57e7 range_7c2854161f56 range_31ca0e513328 range_971ecf19520c range_7726a62825b9 range_a0b0e68be0ef range_197d25f1815d range_83c8e7fbbbd9 range_2551d565c64f range_888c0dcc9a06 range_e6cbfda780d6 range_dc510cebf352 range_1f0be2e860dc range_289b1d4438fc range_1b7049bf2d07 range_d099eddee3b7 range_d39311696efa range_c42451bd4e5e range_d5c4e5d52b13 range_86b2019b7f00 range_3b641ad31af2 range_ebc9206cd758 range_0b2a1f00679b range_46031f4fa053 range_113596af990f range_56e105f89fd9Adds Vitest suites for seed/wechat/diagnostics/QR, many tests updated for WeChat expectations, E2E scripts (`test-channels-stop-start.sh`, `test-model-router-provider-routed-inference.sh`), parity inventory/map updates, scenario/runtime script adjustments (openshell usage, suite selection).range_05442637fae8 range_9b626cae8756 range_594cfcd74cbb range_b8b4fb26d2bf range_5363ee06f9cd range_3fcb1c6114ca range_4f47112eb06c range_37c3af4c9498 range_a54375a6cc47 range_76e4fe287c7a range_57cec7b907c4 range_df8d44ec16de range_54e538eb3e90 range_0096a1ac37c5 range_850ea4a322e1 range_14d3cc9d2039 range_cb160d6d6d94 range_ab6ac293459a range_0af8faba97a0 range_4eb488e7162a range_56e105f89fd9 range_113596af990f range_46031f4fa053 range_2e5966b25a60 range_52828fcaeefd range_77b4d4f26e8c range_e342b6558208 range_0ea8b15d4da8 range_fb4615c399b1 range_828bf8cfb828 range_85245d3e2146 range_3c4a0e04f512 range_4a8a72f08335 range_eeb48b9a7ce8 range_eb55267138e1 range_0e628c745b52 range_6ec16abd3e8a range_dee2a088ca61 range_dc121373ea6f range_5909b101a09f range_212200ad21c2 range_445fe14920a9 range_73dbc269a6ac range_680c5bd46ad3 range_bdd4c51e6f6d range_721f8a0b65c3 range_3683e5dc0e01 range_733539009907 range_b940ebccde50 range_27e0bdbead4b range_c8b512f2091c range_5c0397d64875 range_445e4b17faac range_3276e97664aa range_1baf24dcecf1 range_820a4ca41967 range_3477909ac09c range_f317cecafcde range_053d42eb8df2 range_5de812e10c90 range_151fd9ac19e0 range_bf210b04467d range_ace813ffbc69 range_95396c6ed299 range_deb2cb4dfee7 range_326a80135886 range_ed357bb79970 range_528270ef0ba0 range_625fef322625 range_029a530c9825 range_7b9381929a4e range_1aa4ec167014 range_102004e5ef55 range_f28bde150fd7 range_1e0b683254da range_4a55a603a993 range_4913f8129ff8 range_a9ef29424d50 range_6ff27246355b range_634703111fbf range_bdb832ed1096 range_c79f568a775e range_222567e305d1 range_29959cc67fab range_cf6a3fbf2603 range_c49997ff5666 range_121ca4b6a573 range_a1b2c75d2cc8 range_b65441829d82 range_82ee3e474397 range_50b101d393e3 range_8795054438ae range_9cf47aafb87f range_73c7b111dd2b range_11bc5ec9ec6f range_f014138bcaf4 range_2b20fe3b863b range_63a54cb653d3 range_2ac95e413792 range_a2a593a80d04 range_2b3f46748af0 range_e85df33c8466 range_9ec1691005db range_57d7d0f4cb41 range_c5a504e14441 range_b5485d4eff0a range_22df25d708ca range_5784fc562175 range_5cacb4d05ca8 range_9dfc04788f63 range_d6e82c5c9eab range_bcef0c4c3040 range_eecd2521a093 range_cc4103d44daf range_59a862f59ae4 range_c3942a062bc3 range_eebc37af4966 range_c328076f9ef0 range_400b19748e61 range_f9eefcb96b5e range_e5c7db3f2365 range_c086b1fa73c8 range_005001bdde5b range_bf7c0b544641 range_e1264622aa24 range_2fef70371ec3 range_7446d4c779c6 range_f9bb71ed9244 range_085d59f8a682 range_2b7191adad58 range_f61ce9d0910c range_20ff64cf1ffd range_faa07f3cb690 range_1f58c980a902 range_fc3a8eddf4a9 range_c191310ee617 range_b67663932ec6 range_a72bf50a1782 range_12cb61367505 range_97ffe52a4377 range_87127b1070c0 range_f5df9bce4ae4 range_c135bfb8634a range_6577388d196f range_042a71747c2b range_2dfccd4ec704 range_784421733045 range_b6f936e102d8 range_d41a9a7b6c04 range_848d7636890d✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
⚔️ Resolve merge conflicts
|
| NEMOCLAW_ACCEPT_THIRD_PARTY_SOFTWARE=1 \ | ||
| NEMOCLAW_POLICY_TIER="open" \ | ||
| NEMOCLAW_PROVIDER="routed" \ | ||
| NEMOCLAW_PROVIDER_KEY="${NVIDIA_API_KEY}" \ |
| NEMOCLAW_POLICY_TIER="open" \ | ||
| NEMOCLAW_PROVIDER="routed" \ | ||
| NEMOCLAW_PROVIDER_KEY="${NVIDIA_API_KEY}" \ | ||
| NVIDIA_API_KEY="${NVIDIA_API_KEY}" \ |
| } | ||
|
|
||
| cleanup() { | ||
| local rc=$? |
|
|
||
| cleanup() { | ||
| local rc=$? | ||
| redact_file "$ONBOARD_LOG" |
| cleanup() { | ||
| local rc=$? | ||
| redact_file "$ONBOARD_LOG" | ||
| redact_file "$RESPONSE_LOG" |
| if [ "${NEMOCLAW_E2E_KEEP_SANDBOX:-0}" != "1" ]; then | ||
| nemoclaw "$SANDBOX_NAME" destroy --yes >/dev/null 2>&1 || true | ||
| fi |
| redact_file "$ONBOARD_LOG" | ||
| redact_file "$RESPONSE_LOG" | ||
| redact_file "$HEALTH_LOG" | ||
| if [ "${NEMOCLAW_E2E_KEEP_SANDBOX:-0}" != "1" ]; then |
| redact_file "$RESPONSE_LOG" | ||
| redact_file "$HEALTH_LOG" | ||
| if [ "${NEMOCLAW_E2E_KEEP_SANDBOX:-0}" != "1" ]; then | ||
| nemoclaw "$SANDBOX_NAME" destroy --yes >/dev/null 2>&1 || true |
| redact_file "$RESPONSE_LOG" | ||
| redact_file "$HEALTH_LOG" | ||
| if [ "${NEMOCLAW_E2E_KEEP_SANDBOX:-0}" != "1" ]; then | ||
| nemoclaw "$SANDBOX_NAME" destroy --yes >/dev/null 2>&1 || true |
| if [ "${NEMOCLAW_E2E_KEEP_SANDBOX:-0}" != "1" ]; then | ||
| nemoclaw "$SANDBOX_NAME" destroy --yes >/dev/null 2>&1 || true | ||
| fi | ||
| exit "$rc" |
E2E Advisor RecommendationRequired E2E: None Dispatch hint: Full advisor summaryE2E Recommendation AdvisorBase: Required E2E
Optional E2E
New E2E recommendations
Dispatch hint
|
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/e2e/test-model-router-provider-routed-inference.sh`:
- Around line 129-155: The current test only greps for 'PONG' in the variable
response/RESPONSE_LOG which can yield false positives; update the validation in
the loop and final check to assert both the expected content and evidence it
came from the routed model (e.g., verify the JSON response includes the model
name "nvidia-routed" or a routing metadata field such as "model" or "routed"
alongside the message), by parsing the curl output stored in
response/RESPONSE_LOG and requiring both "PONG" and the routing identifier (from
the HTTP body returned by the https://inference.local/v1/chat/completions call)
before declaring success; ensure the same stricter check replaces the two
grep-only checks that reference response and RESPONSE_LOG so the test only
passes when the reply is from nvidia-routed.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 176a796c-8771-4c9c-8104-27799bb5a548
📒 Files selected for processing (2)
.github/workflows/regression-e2e.yamltest/e2e/test-model-router-provider-routed-inference.sh
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/e2e/test-model-router-provider-routed-inference.sh`:
- Around line 19-37: The helper is_routed_pong_response uses a heredoc (python3
- <<'PY') which overrides stdin so the callers' <<<"$response" never reaches
Python; replace the heredoc invocation with an inline -c invocation that reads
sys.stdin (e.g., python3 -c '...script...' ) or pipe the input into python
(e.g., printf '%s' "$response" | python3 - <<'PY' alternative), keeping the same
Python logic (json.loads(sys.stdin.read()), model/choices parsing, regex check)
so the function reads the caller-provided stdin and returns 0/1 correctly.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 86a1296a-e209-4e25-9cbd-65d6b72159e5
📒 Files selected for processing (1)
test/e2e/test-model-router-provider-routed-inference.sh
| is_routed_pong_response() { | ||
| python3 - <<'PY' | ||
| import json, re, sys | ||
| raw = sys.stdin.read() | ||
| try: | ||
| data = json.loads(raw) | ||
| except Exception: | ||
| raise SystemExit(1) | ||
| model = str(data.get("model", "")) | ||
| choices = data.get("choices") or [] | ||
| content = "" | ||
| if choices and isinstance(choices[0], dict): | ||
| message = choices[0].get("message") or {} | ||
| content = str(message.get("content", "")) | ||
| ok_model = model == "nvidia-routed" or model.startswith("nvidia-routed") | ||
| ok_content = re.search(r"\bPONG\b", content, re.IGNORECASE) is not None | ||
| raise SystemExit(0 if ok_model and ok_content else 1) | ||
| PY | ||
| } |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
payload='{"model":"nvidia-routed","choices":[{"message":{"content":"PONG"}}]}'
broken() {
python3 - <<'PY'
import json, sys
raw = sys.stdin.read()
try:
json.loads(raw)
except Exception:
raise SystemExit(1)
raise SystemExit(0)
PY
}
fixed() {
python3 -c '
import json, sys
raw = sys.stdin.read()
try:
json.loads(raw)
except Exception:
raise SystemExit(1)
raise SystemExit(0)
' <<<"${1-}"
}
if broken <<<"$payload"; then
echo "unexpected: broken version parsed payload"
exit 1
fi
fixed "$payload"
echo "ok: fixed version parses payload from stdin"Repository: NVIDIA/NemoClaw
Length of output: 103
🏁 Script executed:
#!/bin/bash
# Check file structure and headers
echo "=== Checking file structure ==="
wc -l test/e2e/test-model-router-provider-routed-inference.sh
echo -e "\n=== First 10 lines (shebang + SPDX headers) ==="
head -10 test/e2e/test-model-router-provider-routed-inference.sh
echo -e "\n=== Lines 19-37 (is_routed_pong_response function) ==="
sed -n '19,37p' test/e2e/test-model-router-provider-routed-inference.sh
echo -e "\n=== Context around line 159 (first call site) ==="
sed -n '155,165p' test/e2e/test-model-router-provider-routed-inference.sh
echo -e "\n=== Context around line 169 (second call site) ==="
sed -n '165,175p' test/e2e/test-model-router-provider-routed-inference.sh
echo -e "\n=== File permissions (check if executable) ==="
ls -l test/e2e/test-model-router-provider-routed-inference.shRepository: NVIDIA/NemoClaw
Length of output: 2156
🏁 Script executed:
#!/bin/bash
# Check ShellCheck compliance (if available)
if command -v shellcheck &> /dev/null; then
echo "=== Running ShellCheck on the file ==="
shellcheck test/e2e/test-model-router-provider-routed-inference.sh || true
else
echo "ShellCheck not available in sandbox"
fiRepository: NVIDIA/NemoClaw
Length of output: 973
Fix stdin wiring in is_routed_pong_response (currently always fails).
The pattern python3 - <<'PY' makes the heredoc itself stdin for the Python process, so the here-string input <<<"$response" from callers (lines 162 and 172) never reaches sys.stdin.read(). JSON parsing fails every time, breaking the routed completion validation required for #3255 acceptance testing.
Suggested patch
is_routed_pong_response() {
- python3 - <<'PY'
+ python3 -c '
import json, re, sys
raw = sys.stdin.read()
try:
data = json.loads(raw)
except Exception:
raise SystemExit(1)
model = str(data.get("model", ""))
choices = data.get("choices") or []
content = ""
if choices and isinstance(choices[0], dict):
message = choices[0].get("message") or {}
content = str(message.get("content", ""))
ok_model = model == "nvidia-routed" or model.startswith("nvidia-routed")
ok_content = re.search(r"\bPONG\b", content, re.IGNORECASE) is not None
raise SystemExit(0 if ok_model and ok_content else 1)
-PY
+' <<<"${1-}"
}
@@
- if is_routed_pong_response <<<"$response"; then
+ if is_routed_pong_response "$response"; then
@@
-if is_routed_pong_response <<<"$response"; then
+if is_routed_pong_response "$response"; then📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| is_routed_pong_response() { | |
| python3 - <<'PY' | |
| import json, re, sys | |
| raw = sys.stdin.read() | |
| try: | |
| data = json.loads(raw) | |
| except Exception: | |
| raise SystemExit(1) | |
| model = str(data.get("model", "")) | |
| choices = data.get("choices") or [] | |
| content = "" | |
| if choices and isinstance(choices[0], dict): | |
| message = choices[0].get("message") or {} | |
| content = str(message.get("content", "")) | |
| ok_model = model == "nvidia-routed" or model.startswith("nvidia-routed") | |
| ok_content = re.search(r"\bPONG\b", content, re.IGNORECASE) is not None | |
| raise SystemExit(0 if ok_model and ok_content else 1) | |
| PY | |
| } | |
| is_routed_pong_response() { | |
| python3 -c ' | |
| import json, re, sys | |
| raw = sys.stdin.read() | |
| try: | |
| data = json.loads(raw) | |
| except Exception: | |
| raise SystemExit(1) | |
| model = str(data.get("model", "")) | |
| choices = data.get("choices") or [] | |
| content = "" | |
| if choices and isinstance(choices[0], dict): | |
| message = choices[0].get("message") or {} | |
| content = str(message.get("content", "")) | |
| ok_model = model == "nvidia-routed" or model.startswith("nvidia-routed") | |
| ok_content = re.search(r"\bPONG\b", content, re.IGNORECASE) is not None | |
| raise SystemExit(0 if ok_model and ok_content else 1) | |
| ' <<<"${1-}" | |
| } |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@test/e2e/test-model-router-provider-routed-inference.sh` around lines 19 -
37, The helper is_routed_pong_response uses a heredoc (python3 - <<'PY') which
overrides stdin so the callers' <<<"$response" never reaches Python; replace the
heredoc invocation with an inline -c invocation that reads sys.stdin (e.g.,
python3 -c '...script...' ) or pipe the input into python (e.g., printf '%s'
"$response" | python3 - <<'PY' alternative), keeping the same Python logic
(json.loads(sys.stdin.read()), model/choices parsing, regex check) so the
function reads the caller-provided stdin and returns 0/1 correctly.
There was a problem hiding this comment.
Actionable comments posted: 5
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/lib/onboard.ts (1)
5719-5725:⚠️ Potential issue | 🟠 Major | ⚡ Quick winPersist WeChat metadata only after sandbox recreation succeeds.
This writes
current.wechatConfiginto the session before the new sandbox is actually built. If creation fails after this point, the next--resumerun compares against the just-written session value, sohasWechatConfigDrift(session)can be suppressed and the old sandbox can be reused with stale WeChat metadata.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/lib/onboard.ts` around lines 5719 - 5725, The session currently writes current.wechatConfig inside onboardSession.updateSession before the sandbox is recreated, causing stale metadata to be persisted if sandbox creation fails; instead, remove or delay setting current.wechatConfig in that update and only persist the result of toSessionWechatConfig(wechatConfig) after the sandbox recreation step succeeds (i.e., after the sandbox creation function returns successfully), using the same onboardSession.updateSession call to write the wechatConfig; leave telegramConfig and messagingChannelConfig updates as-is and reference onboardSession.updateSession, toSessionWechatConfig, and hasWechatConfigDrift to locate the related logic.
🧹 Nitpick comments (1)
test/e2e/nemoclaw_scenarios/install/dispatch.sh (1)
44-45: ⚡ Quick winAdd a regression test for the
brev-launchablealias.This new alias path is untested in the visible suite; a small dry-run dispatch assertion would prevent silent routing regressions.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@test/e2e/nemoclaw_scenarios/install/dispatch.sh` around lines 44 - 45, Add a regression test that exercises the new alias path "brev-launchable" by invoking the dispatch logic in test/e2e/nemoclaw_scenarios/install/dispatch.sh and asserting a dry-run routing result; update the case branch that currently maps "brev-launchable | launchable)" (and the e2e_install_launchable function) to include a thin dry-run call to the dispatch entry point and an explicit assertion that the resolved route matches the expected launchable handler so the alias is exercised in CI.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/brev-nightly-e2e.yaml:
- Around line 52-53: The job is passing undeclared reusable-workflow inputs
(branch, keep_alive, launchable_id, test_suite, use_launchable,
use_published_launchable) to the called workflow; either remove these parameters
from the call in .github/workflows/brev-nightly-e2e.yaml or add matching
workflow_call input declarations in
.github/workflows/e2e-branch-validation.yaml. Locate the call site where these
keys are supplied (the entries for branch, keep_alive, launchable_id,
test_suite, use_launchable, use_published_launchable) and delete them if they
aren’t needed, or open e2e-branch-validation.yaml and add corresponding inputs
under workflow_call with appropriate defaults/types so the call validates.
In `@src/lib/sandbox/channels.ts`:
- Around line 27-29: channelUsesQrPairing currently infers QR pairing from the
absence of envKey which conflicts with the new loginMethod field (e.g., WeChat
has loginMethod: "host-qr" but an envKey); update the helper
channelUsesQrPairing to prefer the explicit loginMethod (return true when
channel.loginMethod === "host-qr") and fallback to the previous envKey-based
check when loginMethod is undefined so callers correctly detect QR flows for
channels like WeChat; adjust references to channelUsesQrPairing and any call
sites if they assumed the old behavior.
In `@test/e2e/nemoclaw_scenarios/install/launchable.sh`:
- Around line 32-36: The script currently runs rm -rf on clone_dir derived from
E2E_LAUNCHABLE_CLONE_DIR without validation; add a safety guard before removal
to ensure clone_dir is non-empty and not unsafe (e.g., not "/", not ".", not
"$HOME", and not "/tmp" or other critical roots) and fail loudly if validation
fails; use the existing clone_dir/NEMOCLAW_CLONE_DIR variables to perform the
check and only call rm -rf when the path passes validation, otherwise abort with
an error message.
In `@test/e2e/nemoclaw_scenarios/install/repo-current.sh`:
- Around line 26-27: The script currently runs "npm ci --ignore-scripts" without
failing fast; ensure the step aborts the test on failure by making the command's
failure exit the script — either enable errexit at the top (e.g., set -euo
pipefail) or change the line to explicitly check the exit status (e.g., if ! npm
ci --ignore-scripts; then echo "npm ci failed" >&2; exit 1; fi) so downstream
steps cannot continue after a failed "npm ci --ignore-scripts".
In
`@test/e2e/validation_suites/inference/ollama-gpu/01-ollama-chat-completion.sh`:
- Line 37: The current command embeds JSON into a sh -lc string which breaks on
single quotes; replace the sh -lc approach and stream the payload into docker
exec to avoid shell-quoting issues: use echo "$payload" | docker exec -i
"${container_id}" curl -fsS --max-time 30 -H 'Content-Type: application/json' -d
`@-` http://127.0.0.1:11434/v1/chat/completions so curl reads JSON from stdin (no
sh -lc, no fragile interpolation of ${payload} inside the container).
---
Outside diff comments:
In `@src/lib/onboard.ts`:
- Around line 5719-5725: The session currently writes current.wechatConfig
inside onboardSession.updateSession before the sandbox is recreated, causing
stale metadata to be persisted if sandbox creation fails; instead, remove or
delay setting current.wechatConfig in that update and only persist the result of
toSessionWechatConfig(wechatConfig) after the sandbox recreation step succeeds
(i.e., after the sandbox creation function returns successfully), using the same
onboardSession.updateSession call to write the wechatConfig; leave
telegramConfig and messagingChannelConfig updates as-is and reference
onboardSession.updateSession, toSessionWechatConfig, and hasWechatConfigDrift to
locate the related logic.
---
Nitpick comments:
In `@test/e2e/nemoclaw_scenarios/install/dispatch.sh`:
- Around line 44-45: Add a regression test that exercises the new alias path
"brev-launchable" by invoking the dispatch logic in
test/e2e/nemoclaw_scenarios/install/dispatch.sh and asserting a dry-run routing
result; update the case branch that currently maps "brev-launchable |
launchable)" (and the e2e_install_launchable function) to include a thin dry-run
call to the dispatch entry point and an explicit assertion that the resolved
route matches the expected launchable handler so the alias is exercised in CI.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: fc995cc7-c60c-486f-b080-024f91b01676
⛔ Files ignored due to path filters (1)
package-lock.jsonis excluded by!**/package-lock.json
📒 Files selected for processing (97)
.coderabbit.yaml.github/workflows/brev-nightly-e2e.yaml.github/workflows/e2e-scenarios.yaml.github/workflows/nightly-e2e.yaml.github/workflows/regression-e2e.yamlDockerfileDockerfile.baseagents/openclaw/manifest.yamldocs/manage-sandboxes/messaging-channels.mddocs/reference/commands.mdnemoclaw-blueprint/policies/presets/wechat.yamlnemoclaw-blueprint/policies/tiers.yamlnemoclaw-blueprint/scripts/wechat-diagnostics.jspackage.jsonscripts/generate-openclaw-config.pyscripts/nemoclaw-start.shscripts/seed-wechat-accounts.pysrc/ext/wechat/login.test.tssrc/ext/wechat/login.tssrc/ext/wechat/qr.test.tssrc/ext/wechat/qr.tssrc/lib/actions/inference-set.test.tssrc/lib/actions/sandbox/policy-channel.tssrc/lib/actions/sandbox/rebuild.tssrc/lib/adapters/docker/image.tssrc/lib/adapters/docker/index.test.tssrc/lib/agent/defs.test.tssrc/lib/credentials/store.tssrc/lib/host-qr-handlers.tssrc/lib/messaging-channel-config.test.tssrc/lib/messaging-conflict.test.tssrc/lib/messaging-conflict.tssrc/lib/onboard.tssrc/lib/onboard/channel-state.test.tssrc/lib/onboard/channel-state.tssrc/lib/onboard/docker-gpu-patch.test.tssrc/lib/onboard/docker-gpu-patch.tssrc/lib/onboard/dockerfile-patch.test.tssrc/lib/onboard/dockerfile-patch.tssrc/lib/onboard/host-qr-dispatch.tssrc/lib/onboard/initial-policy.tssrc/lib/onboard/messaging-channel-setup.tssrc/lib/onboard/messaging-reuse.test.tssrc/lib/onboard/messaging-reuse.tssrc/lib/onboard/wechat-config.tssrc/lib/policy/index.tssrc/lib/sandbox-base-image.test.tssrc/lib/sandbox-base-image.tssrc/lib/sandbox/build-context.tssrc/lib/sandbox/channels.test.tssrc/lib/sandbox/channels.tssrc/lib/state/onboard-session.test.tssrc/lib/state/onboard-session.tssrc/lib/state/sandbox.tstest/credentials.test.tstest/e2e/docs/parity-inventory.generated.jsontest/e2e/docs/parity-map.yamltest/e2e/nemoclaw_scenarios/expected-states.yamltest/e2e/nemoclaw_scenarios/helpers/emit-context-from-plan.shtest/e2e/nemoclaw_scenarios/install/dispatch.shtest/e2e/nemoclaw_scenarios/install/helpers/install-path-refresh.shtest/e2e/nemoclaw_scenarios/install/launchable.shtest/e2e/nemoclaw_scenarios/install/repo-current.shtest/e2e/nemoclaw_scenarios/onboard/cloud-hermes.shtest/e2e/nemoclaw_scenarios/onboard/cloud-openclaw.shtest/e2e/nemoclaw_scenarios/onboard/local-ollama-openclaw.shtest/e2e/nemoclaw_scenarios/scenarios.yamltest/e2e/runtime/lib/env.shtest/e2e/runtime/run-scenario.shtest/e2e/scenario-framework-tests/e2e-lib-helpers.test.tstest/e2e/scenario-framework-tests/e2e-scenarios-workflow.test.tstest/e2e/test-channels-stop-start.shtest/e2e/test-messaging-providers.shtest/e2e/test-model-router-provider-routed-inference.shtest/e2e/validation_suites/assert/gateway-alive.shtest/e2e/validation_suites/inference/cloud/00-models-health.shtest/e2e/validation_suites/inference/cloud/01-chat-completion.shtest/e2e/validation_suites/inference/cloud/02-inference-local-from-sandbox.shtest/e2e/validation_suites/inference/ollama-auth-proxy/00-proxy-reachable.shtest/e2e/validation_suites/inference/ollama-gpu/00-ollama-models-health.shtest/e2e/validation_suites/inference/ollama-gpu/01-ollama-chat-completion.shtest/e2e/validation_suites/platform/macos/00-macos-smoke.shtest/e2e/validation_suites/sandbox-exec.shtest/e2e/validation_suites/smoke/03-sandbox-shell.shtest/e2e/validation_suites/suites.yamltest/generate-openclaw-config.test.tstest/onboard.test.tstest/policies.test.tstest/policy-tiers.test.tstest/registry.test.tstest/sandbox-build-context.test.tstest/sandbox-provisioning.test.tstest/security-sandbox-tar-traversal.test.tstest/seed-wechat-accounts.test.tstest/snapshot.test.tstest/wechat-diagnostics.test.tsvitest.config.ts
💤 Files with no reviewable changes (1)
- Dockerfile.base
✅ Files skipped from review due to trivial changes (17)
- src/lib/onboard/channel-state.test.ts
- src/lib/agent/defs.test.ts
- src/lib/messaging-channel-config.test.ts
- src/lib/onboard/messaging-reuse.ts
- src/lib/onboard/channel-state.ts
- test/e2e/nemoclaw_scenarios/expected-states.yaml
- package.json
- test/e2e/runtime/lib/env.sh
- src/lib/onboard/messaging-reuse.test.ts
- src/lib/messaging-conflict.ts
- test/sandbox-provisioning.test.ts
- src/lib/credentials/store.ts
- src/lib/onboard/docker-gpu-patch.ts
- scripts/nemoclaw-start.sh
- test/generate-openclaw-config.test.ts
- test/e2e/docs/parity-inventory.generated.json
- agents/openclaw/manifest.yaml
🚧 Files skipped from review as they are similar to previous changes (2)
- .github/workflows/regression-e2e.yaml
- test/e2e/test-model-router-provider-routed-inference.sh
| use_published_launchable: true | ||
| launchable_id: ${{ github.event_name == 'workflow_dispatch' && inputs.launchable_id || '' }} |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/usr/bin/env bash
set -euo pipefail
echo "== Declared workflow_call inputs (.github/workflows/e2e-branch-validation.yaml) =="
awk '
/^[[:space:]]*workflow_call:[[:space:]]*$/ {in_wc=1; next}
in_wc && /^[[:space:]]*inputs:[[:space:]]*$/ {in_inputs=1; next}
in_inputs {
if ($0 ~ /^[[:space:]]{4}[a-zA-Z0-9_-]+:[[:space:]]*$/) {
key=$1; sub(":", "", key); print key
} else if ($0 ~ /^[[:space:]]{2}[a-zA-Z0-9_-]+:[[:space:]]*$/) {
exit
}
}
' .github/workflows/e2e-branch-validation.yaml | sort -u
echo
echo "== Inputs passed by brev-nightly-e2e caller =="
awk '
/^[[:space:]]*with:[[:space:]]*$/ {in_with=1; next}
in_with {
if ($0 ~ /^[[:space:]]{6}[a-zA-Z0-9_-]+:[[:space:]]*/) {
key=$1; sub(":", "", key); print key
} else if ($0 ~ /^[[:space:]]{4}(secrets|permissions|strategy|if|uses|name):[[:space:]]*$/) {
exit
}
}
' .github/workflows/brev-nightly-e2e.yaml | sort -uRepository: NVIDIA/NemoClaw
Length of output: 281
Remove all undeclared reusable-workflow inputs.
The following inputs passed at lines 49-54 are not declared by .github/workflows/e2e-branch-validation.yaml: branch, keep_alive, launchable_id, test_suite, use_launchable, and use_published_launchable. Only secrets is declared as a workflow_call input. Passing undeclared inputs will cause workflow-call validation failure and prevent this job from executing.
🧰 Tools
🪛 actionlint (1.7.12)
[error] 52-52: input "use_published_launchable" is not defined in "./.github/workflows/e2e-branch-validation.yaml" reusable workflow. defined inputs are "branch", "brev_create_timeout_seconds", "brev_gpu_min_vram", "brev_gpu_name", "brev_gpu_type", "brev_provider", "keep_alive", "pr_number", "setup_script_url", "test_suite", "use_launchable"
(workflow-call)
[error] 53-53: input "launchable_id" is not defined in "./.github/workflows/e2e-branch-validation.yaml" reusable workflow. defined inputs are "branch", "brev_create_timeout_seconds", "brev_gpu_min_vram", "brev_gpu_name", "brev_gpu_type", "brev_provider", "keep_alive", "pr_number", "setup_script_url", "test_suite", "use_launchable"
(workflow-call)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/brev-nightly-e2e.yaml around lines 52 - 53, The job is
passing undeclared reusable-workflow inputs (branch, keep_alive, launchable_id,
test_suite, use_launchable, use_published_launchable) to the called workflow;
either remove these parameters from the call in
.github/workflows/brev-nightly-e2e.yaml or add matching workflow_call input
declarations in .github/workflows/e2e-branch-validation.yaml. Locate the call
site where these keys are supplied (the entries for branch, keep_alive,
launchable_id, test_suite, use_launchable, use_published_launchable) and delete
them if they aren’t needed, or open e2e-branch-validation.yaml and add
corresponding inputs under workflow_call with appropriate defaults/types so the
call validates.
| // "host-qr" channels capture the token via a host-side QR handshake instead | ||
| // of a paste prompt. Defaults to "token-paste" when omitted. | ||
| loginMethod?: "token-paste" | "host-qr"; |
There was a problem hiding this comment.
channelUsesQrPairing now disagrees with the new loginMethod contract.
At Line 75, WeChat is marked loginMethod: "host-qr" while still having an envKey (Line 65). But channelUsesQrPairing currently returns !channel.envKey, so it evaluates WeChat as non-QR. Any caller relying on that helper can skip QR flow incorrectly.
💡 Suggested fix
export function channelUsesQrPairing(channel: ChannelDef): boolean {
- return !channel.envKey;
+ return channel.loginMethod === "host-qr" || !channel.envKey;
}Also applies to: 64-76
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/lib/sandbox/channels.ts` around lines 27 - 29, channelUsesQrPairing
currently infers QR pairing from the absence of envKey which conflicts with the
new loginMethod field (e.g., WeChat has loginMethod: "host-qr" but an envKey);
update the helper channelUsesQrPairing to prefer the explicit loginMethod
(return true when channel.loginMethod === "host-qr") and fallback to the
previous envKey-based check when loginMethod is undefined so callers correctly
detect QR flows for channels like WeChat; adjust references to
channelUsesQrPairing and any call sites if they assumed the old behavior.
| local clone_dir="${E2E_LAUNCHABLE_CLONE_DIR:-${HOME}/NemoClaw-launchable-scenario}" | ||
| export NEMOCLAW_CLONE_DIR="${clone_dir}" | ||
| export NEMOCLAW_REF="${NEMOCLAW_REF:-$(git -C "${repo_root}" rev-parse --abbrev-ref HEAD 2>/dev/null || echo HEAD)}" | ||
| rm -rf "${clone_dir}" | ||
| mkdir -p "${clone_dir}" |
There was a problem hiding this comment.
Guard rm -rf against unsafe clone paths.
Line 32-Line 36 delete an env-derived path unconditionally. If E2E_LAUNCHABLE_CLONE_DIR is empty or / (or similarly unsafe), this can wipe unintended directories on the runner.
Suggested hardening
local clone_dir="${E2E_LAUNCHABLE_CLONE_DIR:-${HOME}/NemoClaw-launchable-scenario}"
+ if [ -z "${clone_dir}" ] || [ "${clone_dir}" = "/" ]; then
+ echo "install-launchable: refusing unsafe clone dir: '${clone_dir}'" >&2
+ return 2
+ fi
export NEMOCLAW_CLONE_DIR="${clone_dir}"
export NEMOCLAW_REF="${NEMOCLAW_REF:-$(git -C "${repo_root}" rev-parse --abbrev-ref HEAD 2>/dev/null || echo HEAD)}"
rm -rf "${clone_dir}"📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| local clone_dir="${E2E_LAUNCHABLE_CLONE_DIR:-${HOME}/NemoClaw-launchable-scenario}" | |
| export NEMOCLAW_CLONE_DIR="${clone_dir}" | |
| export NEMOCLAW_REF="${NEMOCLAW_REF:-$(git -C "${repo_root}" rev-parse --abbrev-ref HEAD 2>/dev/null || echo HEAD)}" | |
| rm -rf "${clone_dir}" | |
| mkdir -p "${clone_dir}" | |
| local clone_dir="${E2E_LAUNCHABLE_CLONE_DIR:-${HOME}/NemoClaw-launchable-scenario}" | |
| if [ -z "${clone_dir}" ] || [ "${clone_dir}" = "/" ]; then | |
| echo "install-launchable: refusing unsafe clone dir: '${clone_dir}'" >&2 | |
| return 2 | |
| fi | |
| export NEMOCLAW_CLONE_DIR="${clone_dir}" | |
| export NEMOCLAW_REF="${NEMOCLAW_REF:-$(git -C "${repo_root}" rev-parse --abbrev-ref HEAD 2>/dev/null || echo HEAD)}" | |
| rm -rf "${clone_dir}" | |
| mkdir -p "${clone_dir}" |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@test/e2e/nemoclaw_scenarios/install/launchable.sh` around lines 32 - 36, The
script currently runs rm -rf on clone_dir derived from E2E_LAUNCHABLE_CLONE_DIR
without validation; add a safety guard before removal to ensure clone_dir is
non-empty and not unsafe (e.g., not "/", not ".", not "$HOME", and not "/tmp" or
other critical roots) and fail loudly if validation fails; use the existing
clone_dir/NEMOCLAW_CLONE_DIR variables to perform the check and only call rm -rf
when the path passes validation, otherwise abort with an error message.
| echo "repo-current: npm ci" | ||
| npm ci --ignore-scripts |
There was a problem hiding this comment.
Fail fast when npm ci fails.
Line 27 is not guarded. If the shell is running without errexit, the function can continue after a failed install and produce misleading downstream results.
Suggested fix
echo "repo-current: npm ci"
- npm ci --ignore-scripts
+ if ! npm ci --ignore-scripts; then
+ echo "repo-current: npm ci failed" >&2
+ return 1
+ fi📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| echo "repo-current: npm ci" | |
| npm ci --ignore-scripts | |
| echo "repo-current: npm ci" | |
| if ! npm ci --ignore-scripts; then | |
| echo "repo-current: npm ci failed" >&2 | |
| return 1 | |
| fi |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@test/e2e/nemoclaw_scenarios/install/repo-current.sh` around lines 26 - 27,
The script currently runs "npm ci --ignore-scripts" without failing fast; ensure
the step aborts the test on failure by making the command's failure exit the
script — either enable errexit at the top (e.g., set -euo pipefail) or change
the line to explicitly check the exit status (e.g., if ! npm ci
--ignore-scripts; then echo "npm ci failed" >&2; exit 1; fi) so downstream steps
cannot continue after a failed "npm ci --ignore-scripts".
| # Docker GPU host networking gives the sandbox a direct loopback path to | ||
| # Ollama; use docker exec like legacy test-gpu-e2e.sh instead of the normal | ||
| # OpenShell dashboard/gateway forward path. | ||
| body="$(docker exec "${container_id}" sh -lc "curl -fsS --max-time 30 -H 'Content-Type: application/json' -d '$payload' http://127.0.0.1:11434/v1/chat/completions")" |
There was a problem hiding this comment.
Avoid shell-embedding JSON payload in docker exec command.
Quoting ${payload} inside sh -lc is brittle and can break on single quotes in model names. Pass curl args directly to docker exec instead.
Proposed fix
-body="$(docker exec "${container_id}" sh -lc "curl -fsS --max-time 30 -H 'Content-Type: application/json' -d '$payload' http://127.0.0.1:11434/v1/chat/completions")"
+body="$(docker exec "${container_id}" \
+ curl -fsS --max-time 30 \
+ -H "Content-Type: application/json" \
+ -d "${payload}" \
+ "http://127.0.0.1:11434/v1/chat/completions")"📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| body="$(docker exec "${container_id}" sh -lc "curl -fsS --max-time 30 -H 'Content-Type: application/json' -d '$payload' http://127.0.0.1:11434/v1/chat/completions")" | |
| body="$(docker exec "${container_id}" \ | |
| curl -fsS --max-time 30 \ | |
| -H "Content-Type: application/json" \ | |
| -d "${payload}" \ | |
| "http://127.0.0.1:11434/v1/chat/completions")" |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@test/e2e/validation_suites/inference/ollama-gpu/01-ollama-chat-completion.sh`
at line 37, The current command embeds JSON into a sh -lc string which breaks on
single quotes; replace the sh -lc approach and stream the payload into docker
exec to avoid shell-quoting issues: use echo "$payload" | docker exec -i
"${container_id}" curl -fsS --max-time 30 -H 'Content-Type: application/json' -d
`@-` http://127.0.0.1:11434/v1/chat/completions so curl reads JSON from stdin (no
sh -lc, no fragile interpolation of ${payload} inside the container).
|
Closing in favor of rebased branch |
## Summary Adds a regression E2E guard for model-router provider-routed inference so the routed inference path is covered in the regression workflow. This replaces closed PR #3594 from the rebased branch because repository rules blocked force-pushing the original PR branch. ## Related Issue Refs #3255 ## Changes - Adds `test/e2e/test-model-router-provider-routed-inference.sh` for provider-routed inference coverage. - Updates `.github/workflows/regression-e2e.yaml` to include the new regression guard. - Updates E2E parity metadata in `test/e2e/docs/parity-map.yaml` and `test/e2e/docs/parity-inventory.generated.json`. - Carries the rebased compatibility cleanup currently present on `pr-3594-rebased-main`. ## Type of Change - [x] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Verification - [ ] `npx prek run --all-files` passes - [ ] `npm test` passes - [x] Tests added or updated for new or changed behavior - [x] No secrets, API keys, or credentials committed - [ ] Docs updated for user-facing behavior changes - [ ] `make docs` builds without warnings (doc changes only) - [ ] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) --- <!-- DCO sign-off required by CI. Run: git config user.name && git config user.email --> Signed-off-by: Julie Yaunches <jyaunches@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added an end-to-end coverage guard that verifies provider-routed Model Router onboarding and routed completions, including health checks, routed response validation, redacted log capture, and failure reporting. * **Chores** * CI regression workflow updated to include a selectable job for the new E2E guard and conditionally run it; model-router-specific artifacts are uploaded on failures. * **Documentation** * Test inventory and mapping updated to include the new script, assertions, and revised totals. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/NemoClaw/pull/3601) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary Fixes #3255 by correcting the Model Router pool config used by the Provider Routed path: - routes `nvapi-*` keys to `https://integrate.api.nvidia.com/v1` - fixes the Nemotron 3 Nano LiteLLM model ID - fixes the Nemotron 3 Super LiteLLM model ID - adds unit regression coverage so these config values do not drift back ## Validation - `npm test -- --run test/validate-blueprint.test.ts` - `npm test -- --run test/onboard.test.ts -t "Model Router"` - `npm run build:cli` ## E2E guard Failing-test-first guard: #3594 RED evidence on main-equivalent code: https://github.com/NVIDIA/NemoClaw/actions/runs/25922557128 After this PR is open, dispatch: ```bash gh workflow run regression-e2e.yaml \ --repo NVIDIA/NemoClaw \ -f jobs=model-router-provider-routed-inference-e2e \ --ref issue-3255-model-router-provider-routed-503 ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated Nemotron model configurations with revised identifiers and API endpoint settings. * **Tests** * Added validation tests to ensure model routing and identifier configurations function correctly. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/NemoClaw/pull/3614) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai -->
Summary
Adds a failing-test-first regression guard for #3255.
The new
model-router-provider-routed-inference-e2ejob lives in.github/workflows/regression-e2e.yamlonly. It is not scheduled nightly and should be dispatched manually while the fix is in flight.Expected behavior on unfixed main-equivalent code
This PR is expected to fail until #3255 is fixed. The guard onboards with
NEMOCLAW_PROVIDER=routed, then asserts:model-routerreports at least one healthy endpoint, andhttps://inference.local/v1/chat/completionsreturns a routednvidia-routedcompletion.Expected RED fragment on unfixed code:
or:
Dispatch
Notes
NVIDIA_API_KEYsecret.Related: #3255
Summary by CodeRabbit