Skip to content

fix(reasoning): hide unsupported effort choices in model pickers - #85246

Closed
fangliquanflq wants to merge 7 commits into
NousResearch:mainfrom
fangliquanflq:fix/reasoning-effort-model-support
Closed

fangliquanflq wants to merge 7 commits into
NousResearch:mainfrom
fangliquanflq:fix/reasoning-effort-model-support

Conversation

@fangliquanflq

Copy link
Copy Markdown

What does this PR do?

Model pickers in Desktop and Dashboard currently offer the full reasoning-effort scale for every reasoning model, including levels the selected model cannot represent. This change preserves exact models.dev reasoning controls in the shared model-options payload and filters both pickers, preventing users from saving unsupported choices while keeping the existing full-list fallback for models without metadata.

Symptom

Selecting a model with a restricted reasoning scale, such as a model that supports only low, high, and max, still shows minimal, medium, xhigh, and ultra. The Off choice also appears even when the model exposes effort selection without toggle support.

Impact

Desktop and Dashboard users can select reasoning values that the active model does not support. Those choices are then folded or ignored downstream, so the picker advertises control that the model cannot honor.

Bug Cause

Trigger: hermes_cli/inventory.py:404 / _apply_capabilities, followed by the Desktop and Dashboard reasoning pickers.

Causal chain:

  1. models.dev supplies exact reasoning_options for a model.
  2. The shared model-options payload reduces that metadata to one reasoning boolean.
  3. Both UI surfaces render their static canonical option lists and expose unsupported values.

Why it is wrong: The backend discards the model-specific effort and toggle capabilities before either UI can use them.

Working sibling / contrast: Models without known metadata safely use the existing complete effort list. The new filtering applies only when exact capability metadata is available.

Ruled out: Provider request serialization is not the source of the picker mismatch. The unsupported choices are already visible before a request is sent, and tests reproduce the defect entirely in metadata parsing, payload shaping, and UI option filtering.

Fix

Parse models.dev effort and toggle options into shared model capabilities, expose them through /api/model/options, and consume the same metadata in Desktop and Dashboard. The UI helpers preserve canonical ordering, omit Off when toggling is unsupported, select a supported fallback when a saved/default effort is invalid, and retain the full scale for unknown or older metadata.

Related Issue

Closes #85209

Type of Change

  • Bug fix (non-breaking change that fixes an issue)

Changes Made

  • agent/models_dev.py - preserve exact reasoning efforts and toggle support from models.dev metadata.
  • hermes_cli/inventory.py and hermes_cli/web_server.py - include model-specific reasoning controls in the shared model-options payload.
  • apps/desktop/src/ - filter model effort rows and toggle behavior with supported fallbacks.
  • web/src/ - filter Dashboard reasoning options from the same capability payload.
  • tests/, apps/desktop/src/**/*.test.ts*, and web/src/**/*.test.ts - cover metadata parsing, payload propagation, fallback compatibility, and both UI consumers.

How to Test

  1. Open a model picker for a model whose metadata declares a restricted effort set and verify only those levels appear.
  2. Verify Off appears only when the model supports disabling reasoning, and that models without exact metadata retain the complete list.
  3. Run the related automated tests:
scripts/run_tests.sh tests/agent/test_models_dev.py tests/hermes_cli/test_inventory.py
cd web && npm exec vitest -- run src/lib/reasoning-effort.test.ts && npm run typecheck
cd apps/desktop && npx vitest run src/lib/reasoning-effort.test.ts src/app/shell/model-edit-submenu.test.tsx

Verified results: 36 Python tests passed, 9 Dashboard tests passed with Dashboard typecheck, and 15 Desktop tests passed.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run the repository test entry on the relevant Python tests and all tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: Windows 11

Documentation & Housekeeping

  • Relevant documentation update: N/A - no user-facing configuration or workflow changed
  • cli-config.yaml.example update: N/A - no config keys changed
  • CONTRIBUTING.md or AGENTS.md update: N/A - no architecture or workflow changed
  • I've considered cross-platform impact per the compatibility guide - the changed logic is platform-independent
  • Tool descriptions/schemas update: N/A - no model tool behavior changed

Screenshots / Logs

Automated behavior tests cover the exact supported-option sets, toggle visibility, supported fallback selection, and unknown-metadata compatibility.

@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/desktop Electron desktop app (apps/desktop/*) comp/dashboard Web dashboard / control panel UI (dashboard/, landing) comp/tools Tool registry, model_tools, toolsets P3 Low — cosmetic, nice to have labels Aug 13, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference, author can ignore or act on any point.

fix(reasoning): hide unsupported effort choices in model pickers

  1. agent/tool_executor.py (_run_sequential_tool_execution_middleware): applying HERMES_CONCURRENT_TOOL_TIMEOUT_S (default 420s) to all sequential tools is a behavior change — sequential dispatch previously had no deadline. Long-running tools not in _NEVER_PARALLEL_TOOLS (e.g. terminal running a build, execute_code doing heavy compute) will now be force-timed-out mid-execution and reported to the model as failures even if they complete. Please confirm the exclusion set covers every legitimately long-running tool, or document the new ceiling.

  2. After a timeout, the abandoned daemon worker keeps executing the tool (threads cannot be cancelled). The design correctly suppresses the late post_tool_call hook and marks the result effect_disposition="unknown", but the tool's side effects (file writes, message sends, state mutations) can still land after the timeout message was already surfaced to the model. A comment noting this inherent limitation would help future readers not assume cancellation.

  3. agent/models_dev.py (get_model_capabilities): reasoning_efforts preserves registry order (deduped via dict.fromkeys), while both UI helpers (supportedReasoningEfforts desktop, effortOptionsForModel web) re-sort to canonical order. Consumers normalize, so behavior is correct — normalizing once in models_dev would make the API contract deterministic.

  4. reasoning_toggle conflates a real {"type":"toggle"} option with a "none" value inside an effort list. Both map to "Off is supported", which is what both UIs need, so behavior is correct — just noting the semantic is broader than the name implies.

  5. Nit: web/vite.config.test.ts (in the sibling provenance PR) notwithstanding, web/src/lib/reasoning-effort.ts effortOptionsForModel returns EFFORT_OPTIONS by reference in the unknown-metadata fallback (return EFFORT_OPTIONS) — if any caller ever mutated the result it would corrupt the module constant. The .filter paths return fresh arrays, so only the fallback is shared; consider [...EFFORT_OPTIONS] for safety.

@teknium1

Copy link
Copy Markdown
Collaborator

Thanks @fangliquanflq — the direction is sound in principle, but I live-tested the premise before salvaging and the catalog data can't support it.

Sent every ladder level (none minimal low medium high xhigh max) with a 1-token prompt to 7 OpenRouter models (49 calls) and compared against each model's published supported_efforts:

model catalog list unlisted levels → result
anthropic/claude-sonnet-4.6 max/high/medium/low none, minimal, xhigh → all 200
openai/gpt-5.4 xhigh/high/medium/low/none minimal → 200 (rt=11); max → 200 (rt=45, the most reasoning tokens of any level)
z-ai/glm-5.3 max/high/low minimal, medium, xhigh → 200; only none 400s (mandatory)
x-ai/grok-4.6 xhigh/high/medium/low minimal, max → 200; only none 400s (mandatory)
deepseek/deepseek-v4-pro xhigh/high none/minimal/low/medium/max → all 200
moonshotai/kimi-k3 max/high/low minimal, medium, xhigh → all 200
qwen/qwen3.6-plus (none published) every level 200

No unlisted level was rejected anywhere; the only 400s were "reasoning is mandatory" on none, which main already handles (can_disable_reasoning from the catalog's mandatory flag, #90412 — the Desktop hides the Off toggle on it). Aggregators translate effort names upstream rather than reject them, so supported_efforts (and models.dev's reasoning_options) describe the vendor's native vocabulary, not what the route accepts. Filtering the picker by those lists would hide levels that demonstrably work, including the strongest one on gpt-5.4. That was also the standing ruling in #90412 ("supported_efforts is deliberately NOT forwarded").

Closing on that evidence rather than merit. If a vendor's direct endpoint (not an aggregator) ever rejects unlisted levels, a per-provider send-path clamp (agent/reasoning_effort.py::clamp_effort, the shape used for Codex/Kimi) is the right fix, not picker filtering.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard comp/dashboard Web dashboard / control panel UI (dashboard/, landing) comp/desktop Electron desktop app (apps/desktop/*) comp/tools Tool registry, model_tools, toolsets P3 Low — cosmetic, nice to have type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Model picker should show only the reasoning-effort levels a model actually supports

4 participants