Skip to content

fix(reasoning): normalize custom-provider model ids for fallback detection - #3327

Closed
Carry00 wants to merge 1 commit into
nesquena:masterfrom
Carry00:fix/custom-provider-reasoning-normalization
Closed

Carry00 wants to merge 1 commit into
nesquena:masterfrom
Carry00:fix/custom-provider-reasoning-normalization

Conversation

@Carry00

@Carry00 Carry00 commented Jun 1, 2026

Copy link
Copy Markdown
Contributor

Thinking Path

  • Hermes WebUI should expose the reasoning-effort selector for thinking-capable models even when they are served through named custom:* providers.
  • The current fallback path already handles slash-prefixed ids and a small set of hand-written bare-name prefixes.
  • That approach fixed deepseek-v4-flash, but it still depends on exact separator shape and misses equivalent ids such as deepseek.v3.2 or deepseek_v4_flash.
  • Custom API aggregators frequently rewrite model ids with dots, underscores, or extra vendor namespaces, so format-specific prefix checks are inherently brittle.
  • This PR makes the fallback heuristic normalize non-slash model ids before matching them against reasoning-capable model families.
  • The result is a more durable fix for custom-provider model naming drift, while also narrowing a known false-positive path in the old substring-based keyword fallback.

What Changed

  • Added _reasoning_name_candidates() in api/config.py to generate normalized candidate names for heuristic reasoning-capability checks.
  • Added _candidate_supports_reasoning() to match reasoning-capable model families by normalized tokens/family semantics instead of raw string shape.
  • Replaced the previous custom-provider fallback block in _heuristic_reasoning_efforts() with candidate-based family matching.
  • Kept the existing resolution order unchanged:
    • provider-specific authoritative paths first
    • models.dev metadata next
    • heuristic fallback last
  • Added regression coverage for separator and namespace variants:
    • deepseek.v3.2
    • deepseek_v3_2
    • vendor.deepseek.v3.2
    • deepseek.v4-flash
    • deepseek_v4_flash
  • Added negative coverage to prevent vendor-prefix keyword false positives:
    • thinkinghub.llama-3.1-70b
    • reasoninghub.llama-3.1-70b
  • Updated CHANGELOG.md with the user-visible behavior change.

Why It Matters

  • Named custom providers do not consistently preserve model-id separator style.
  • The old heuristic assumed a few exact shapes (vendor/model, vendor.model, selected bare prefixes), so equivalent DeepSeek ids could silently hide the reasoning-effort selector.
  • This PR improves resilience to naming variation without broadening the authoritative paths or changing built-in provider behavior.
  • It also fixes an adjacent correctness issue where substring-based keyword matching could falsely enable reasoning for unrelated vendor prefixes like thinkinghub.*.

Verification

  • Runtime verification via direct resolve_model_reasoning_efforts() calls:
    • deepseek.v3.2 -> returns full supported efforts
    • deepseek_v3_2 -> returns full supported efforts
    • vendor.deepseek.v3.2 -> returns full supported efforts
    • thinkinghub.llama-3.1-70b -> returns []
    • reasoninghub.llama-3.1-70b -> returns []
  • python3 -m py_compile api/config.py
  • VS Code diagnostics checked for:
    • api/config.py
    • tests/test_custom_provider_bare_model_reasoning.py
    • CHANGELOG.md

Risks / Follow-ups

  • This remains a heuristic fallback for cases where no authoritative metadata exists, so future provider-specific naming schemes may still need family-rule expansion.
  • A stronger long-term direction would be allowing explicit capability overrides for named custom providers/models, but that is intentionally out of scope here.
  • Automated pytest execution could not be completed in this environment because pytest is not installed (python3 -m pytest -> No module named pytest).

Model Used

  • GPT-5.4 via Trae coding agent
  • Tooling used: repository search/read, patch editing, local Python runtime checks, diagnostics

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Shipped in v0.51.198 (stage-batch10, via #3350). Thanks @Carry00 — the token-aware normalization (vendor-namespace stripping + separator folding) cleanly fixes the reasoning-effort detection for custom:* providers. Opus + Codex both confirmed it only affects the custom-provider fallback heuristic path and doesn't regress existing slash/@-qualified ids. 🎉

AJV20 pushed a commit to AJV20/hermes-webui that referenced this pull request Jun 1, 2026
…lice doc + nesquena#3341 profile skill counts

nesquena#3327 fix(reasoning): normalize custom-provider model ids for fallback heuristics
Co-authored-by: Carry00 <Carry00@users.noreply.github.com>

nesquena#3334 docs(rfc): mark run-adapter Slice 4f shipped, define Slice 4g gate
Co-authored-by: Michaelyklam <Michaelyklam@users.noreply.github.com>

nesquena#3341 fix(profiles): show enabled vs compatible skill counts
Co-authored-by: b3nw <b3nw@users.noreply.github.com>
AJV20 pushed a commit to AJV20/hermes-webui that referenced this pull request Jun 1, 2026
eleboucher pushed a commit to eleboucher/homelab that referenced this pull request Jun 2, 2026
…➔ 0.51.210) (#782)

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
| [ghcr.io/nesquena/hermes-webui](https://github.com/nesquena/hermes-webui) | patch | `0.51.197` → `0.51.210` |

---

### Release Notes

<details>
<summary>nesquena/hermes-webui (ghcr.io/nesquena/hermes-webui)</summary>

### [`v0.51.210`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051210--2026-06-02--Release-GD-stage-batch1--model-picker-multi-slash-fix--extensionless-preview-highlighting)

[Compare Source](nesquena/hermes-webui@v0.51.209...v0.51.210)

##### Fixed

- Model picker no longer snaps to the wrong model when multiple multi-slash model IDs from the same proxy provider share the same base name. Exact-match priority in `_findModelInDropdown` and first-segment-only stripping in `_normalizeConfiguredModelKey` / `_norm_model_id` prevent collisions in selection, badge assignment, and configured-entry dedup ([#&#8203;3360](nesquena/hermes-webui#3360), [@&#8203;b3nw](https://github.com/b3nw)).
- Workspace file previews now syntax-highlight common code/config filenames without useful extensions, including `Dockerfile`, `Dockerfile.*`, `Makefile`, `GNUmakefile`, `CMakeLists.txt`, `.gitignore`, and `.dockerignore` ([#&#8203;3365](nesquena/hermes-webui#3365), [@&#8203;AJV20](https://github.com/AJV20)).

### [`v0.51.209`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051209--2026-06-02--Release-GC-WebUI-dashboard-plugin-system-with-iframe-isolation)

[Compare Source](nesquena/hermes-webui@v0.51.208...v0.51.209)

##### Added

- WebUI dashboard plugins: plugins that ship a UI under `~/.hermes/plugins/<name>/dashboard/` (with a `manifest.json`) now appear as opt-in cards in Settings → Plugins (default off). Once enabled, an **Open** button renders the plugin page inside a sandboxed iframe (`sandbox="allow-scripts allow-forms allow-popups"` — no `allow-same-origin`, so plugin JS/CSS/modals stay fully isolated from the parent app). New `/plugins/` (shared assets) and `/dashboard-plugins/<name>/` (per-plugin assets) static routes serve only built `dist/`/`static/` files with path-traversal, dotfile, and extension-allowlist protection (plugin source/config such as `plugin_api.py`/`manifest.json`/`.env` is never served), and both the page and asset routes are gated server-side on the enable state + an HTTP `sandbox` CSP + `nosniff`. Plugin `name` and `tab.path` are validated at load. Display-only — no plugin backend/subprocess execution ([#&#8203;2622](nesquena/hermes-webui#2622), [@&#8203;pix0127](https://github.com/pix0127)).

### [`v0.51.208`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051208--2026-06-02--Release-GB-workspace-upload-hardening-hotfix)

[Compare Source](nesquena/hermes-webui@v0.51.207...v0.51.208)

##### Fixed

- Hardened the workspace file-upload surface ([#&#8203;3104](nesquena/hermes-webui#3104) follow-up): (1) a negative `Content-Length` no longer bypasses the size cap and triggers an unbounded `rfile.read(-1)` — the length is now validated `[0, MAX_UPLOAD_BYTES]` centrally in `parse_multipart` for every upload handler; (2) `.tar`, `.tbz2`, and `.txz` archives now auto-extract (the upload handler's archive-suffix set was narrower than `extract_archive`'s, so those silently landed as raw files); (3) a rejected archive (zip-slip / zip-bomb / corrupt / too-many-members) now surfaces an error toast in the workspace panel instead of a misleading "Uploaded" success; (4) an in-workspace symlink subpath can no longer make the upload target `mkdir`/write outside the workspace root. Regression tests added.

### [`v0.51.207`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051207--2026-06-02--Release-GA-Edge-TTS-as-an-alternative-speech-engine)

[Compare Source](nesquena/hermes-webui@v0.51.206...v0.51.207)

##### Added

- Added an optional server-side **Edge TTS** speech engine (Microsoft neural voices) selectable in Settings → Preferences → TTS Engine, alongside the existing browser speech synthesis. The voice list switches to the Edge neural voices when selected. A new `POST /api/tts` endpoint streams the audio, gated by the same-origin CSRF check + session auth, a per-client rate limit, a 5000-character cap, and a voice allowlist. `edge-tts` is an optional dependency — the endpoint returns a clear install hint (503) when it isn't present, so existing installs are unaffected ([#&#8203;2931](nesquena/hermes-webui#2931), [@&#8203;liuqiangweb-svg](https://github.com/liuqiangweb-svg)).

### [`v0.51.206`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051206--2026-06-02--Release-FZ-workspace-file-upload--drag-and-drop-with-archive-extraction)

[Compare Source](nesquena/hermes-webui@v0.51.205...v0.51.206)

##### Added

- Workspace file panel: an **Upload** button and drag-and-drop that POST to a new `/api/workspace/upload` endpoint. Files land in the session workspace (resolved via the trusted-workspace guard), are de-duplicated with `-1`/`-2` suffixes, and archives (`.zip`/`.tar.*`) are auto-extracted into the target subdirectory with zip-bomb (size-cap + member-count-cap) and zip-slip (path-containment) protections. The extraction size cap is tunable via `HERMES_WEBUI_MAX_EXTRACTED_MB` (defaults to 10× the upload cap). Extraction errors are surfaced to the frontend instead of being silently swallowed, and the archive is removed on failure ([#&#8203;3104](nesquena/hermes-webui#3104), [@&#8203;antoniocarlos97ss](https://github.com/antoniocarlos97ss)).

### [`v0.51.205`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051205--2026-06-01--Release-FY-stage-hi1--workspace-syntax-highlighting--generated-image-cards--manual-title-regeneration)

[Compare Source](nesquena/hermes-webui@v0.51.204...v0.51.205)

##### Added

- Workspace file previews now render with syntax highlighting via Prism.js (already loaded for chat code blocks), covering common languages (Python, JS/TS, CSS, JSON, SQL, shell, and more) and degrading gracefully to plain text for unknown/plain files and when offline. The preview code surface uses a single uniform background across light and dark themes ([#&#8203;3337](nesquena/hermes-webui#3337), [@&#8203;mysoul12138](https://github.com/mysoul12138)).
- Generated local image artifacts now render as a clean inline image (with click-to-zoom lightbox) plus a hover/focus-revealed **Download** action overlaid on the image, served through authenticated `/api/media` URLs — matching the common AI-chat pattern of letting the image be the hero rather than wrapping it in a permanent card ([#&#8203;3220](nesquena/hermes-webui#3220), [@&#8203;AJV20](https://github.com/AJV20)).
- The session action menu can regenerate conversation titles on demand from the saved transcript, updating the sidebar without touching conversation chronology and syncing the new title through to state.db when Insights sync is enabled. The menu was also streamlined to a compact icon + label layout (descriptions move to hover tooltips). Closes [#&#8203;3106](nesquena/hermes-webui#3106) ([#&#8203;3223](nesquena/hermes-webui#3223), [@&#8203;AJV20](https://github.com/AJV20)).

### [`v0.51.204`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051204--2026-06-01--Release-FX-stage-batch17--projectsession-operations-honor-the-sessions-own-profile)

[Compare Source](nesquena/hermes-webui@v0.51.203...v0.51.204)

##### Fixed

- Project and session operations (project create/rename/recolor/delete/unassign, session move, and the profile chip label) now key on the session's own profile (`S.session.profile`) instead of the global active profile, so switching between sessions from different profiles no longer causes silent 404s, misleading chip labels, or project-picker entries from the wrong profile. The project picker also filters to the session's profile and surfaces an error toast on failure instead of a silent no-op ([#&#8203;3331](nesquena/hermes-webui#3331), [@&#8203;PINKIIILQWQ](https://github.com/PINKIIILQWQ)).

### [`v0.51.203`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051203--2026-06-01--Release-FW-stage-batch15--sticky-manual-unpin-for-streaming-chat-scroll)

[Compare Source](nesquena/hermes-webui@v0.51.202...v0.51.203)

##### Changed

- Streaming chat scroll now uses a sticky manual-unpin model: once you scroll up to read earlier content during a streaming response, the view stays put and no longer auto-follows the live tail until you scroll back to the bottom (near-bottom hysteresis on downward motion) or click the scroll-to-bottom control. Tool cards, token updates, and layout growth no longer re-pin the viewport after a reading pause. This replaces the [#&#8203;3250](nesquena/hermes-webui#3250) upward-intent timeout and supersedes the v0.51.199 proximity-re-pin ([#&#8203;3330](nesquena/hermes-webui#3330)), matching the streaming-scroll behavior of ChatGPT/Claude/Codex. Fresh streams reset the follow state on attach ([#&#8203;3343](nesquena/hermes-webui#3343), [@&#8203;pamnard](https://github.com/pamnard)).

### [`v0.51.202`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051202--2026-06-01--Release-FV-stage-batch14--filter-interrupted-recovery-control-text-from-visible-transcript)

[Compare Source](nesquena/hermes-webui@v0.51.201...v0.51.202)

##### Fixed

- Interrupted SSE-recovery control text (the synthetic `stale_interrupted_event` run-journal payload) is now kept out of the visible chat transcript instead of being replayed as a message: it's marked `recovery_control` on the backend and filtered across the `msgContent()` render path, the SSE settle/error handlers, and final transcript filtering, so platform-only control state no longer leaks into the conversation ([#&#8203;3321](nesquena/hermes-webui#3321), [@&#8203;franksong2702](https://github.com/franksong2702)).

### [`v0.51.201`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051201--2026-06-01--Release-FU-stage-batch13--colored-diff-lines-in-tool-card-snippets)

[Compare Source](nesquena/hermes-webui@v0.51.200...v0.51.201)

##### Added

- Tool-card result snippets that contain a unified diff now render with the same green/red/cyan diff coloring already used for diffs in chat messages (reusing the existing `.diff-block` styles), with an expand/collapse toggle that preserves the coloring. Non-diff snippets are unchanged ([#&#8203;3336](nesquena/hermes-webui#3336), [@&#8203;mysoul12138](https://github.com/mysoul12138)).

### [`v0.51.200`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051200--2026-06-01--Release-FT-stage-batch12--remote-gateway-health-probe--ephemeral-turn-field-preservation)

[Compare Source](nesquena/hermes-webui@v0.51.199...v0.51.200)

##### Fixed

- The Tasks/Cron panel no longer shows a spurious "Gateway not configured" banner in multi-container Docker deployments where the WebUI image doesn't ship the `gateway` Python package: agent-health now probes the remote gateway via `HERMES_API_URL` before falling back to the local `gateway.status` import. Closes [#&#8203;3281](nesquena/hermes-webui#3281) ([#&#8203;3312](nesquena/hermes-webui#3312), [@&#8203;Sanjays2402](https://github.com/Sanjays2402)).
- Force-reloading the active session (`loadSession(sid, {forceReload:true})`) no longer drops ephemeral turn fields (`_turnUsage`, `_turnDuration`, `_turnTps`, `_gatewayRouting`, `_statusCard`): the ephemeral-field carry-forward now reads the prior `S.messages` before it's reset, so the token-usage badge and status cards survive an external refresh. Closes [#&#8203;3306](nesquena/hermes-webui#3306) ([#&#8203;3313](nesquena/hermes-webui#3313), [@&#8203;Sanjays2402](https://github.com/Sanjays2402)).

### [`v0.51.199`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051199--2026-06-01--Release-FS-stage-batch11--pinned-scroll-recovery--inline-math-currency-false-positive)

[Compare Source](nesquena/hermes-webui@v0.51.198...v0.51.199)

##### Fixed

- Pinned chat now recovers its scroll position after a DOM rebuild: `_setMessageScrollToBottom` retries on the next layout frame, and `scrollIfPinned` re-pins when the pane has drifted more than 500px from the bottom, so a message-list rebuild no longer leaves a pinned conversation stranded mid-scroll. Closes [#&#8203;3319](nesquena/hermes-webui#3319) ([#&#8203;3330](nesquena/hermes-webui#3330), [@&#8203;jianongHe](https://github.com/jianongHe)).
- The `$...$` inline-math renderer no longer treats currency like `$1,000 xuống ~$95` as math: the opening `$` followed by a digit is now rejected (aligning with smd's `se()` guard), so dollar amounts render as plain text. Digit-leading inline math (e.g. `$2x = 4$`) should now use the LaTeX-style `\(2x = 4\)` or display `$$2x = 4$$` delimiters ([#&#8203;3311](nesquena/hermes-webui#3311), [@&#8203;toanalien](https://github.com/toanalien)).

### [`v0.51.198`](https://github.com/nesquena/hermes-webui/blob/HEAD/CHANGELOG.md#v051198--2026-06-01--Release-FR-stage-batch10--custom-provider-reasoning-model-id-normalize--profile-skill-counts--run-adapter-RFC-slice)

[Compare Source](nesquena/hermes-webui@v0.51.197...v0.51.198)

##### Fixed

- Reasoning-effort detection for named `custom:*` providers now normalizes non-slash model ids before applying its fallback family heuristics, so separator variants such as `deepseek.v3.2`, `deepseek_v4_flash`, and vendor-namespaced ids like `vendor.deepseek.v3.2` resolve the same way as `deepseek-v4-flash`. The keyword fallback is now token-aware rather than substring-based, preserving names like `model-thinking-preview` without falsely enabling reasoning for unrelated prefixes such as `thinkinghub.llama-3.1-70b` ([#&#8203;3327](nesquena/hermes-webui#3327), [@&#8203;Carry00](https://github.com/Carry00)).
- Profile cards now show enabled vs compatible skill counts (computed with an 8s TTL cache that clears on profile switch) instead of a single ambiguous count. Closes [#&#8203;3339](nesquena/hermes-webui#3339) ([#&#8203;3341](nesquena/hermes-webui#3341), [@&#8203;b3nw](https://github.com/b3nw)).

##### Changed

- The [#&#8203;1925](nesquena/hermes-webui#1925) runtime-adapter RFC now marks the configured runner-client boundary as shipped in v0.51.188 ([#&#8203;3073](nesquena/hermes-webui#3073) / [#&#8203;3274](nesquena/hermes-webui#3274)) and defines the next Slice 4g gate for a supervised local runner process harness: real runner-owned `AIAgent` execution, restart/reattach proof, bounded runner health diagnostics, and no new WebUI runtime-surrogate globals ([#&#8203;3334](nesquena/hermes-webui#3334), [@&#8203;Michaelyklam](https://github.com/Michaelyklam)).

</details>

---

### Configuration

📅 **Schedule**: Branch creation - At any time (no schedule defined), Automerge - At any time (no schedule defined).

🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied.

♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 **Ignore**: Close this PR and you won't be reminded about these updates again.

---

 - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box

---

This PR has been generated by [Renovate Bot](https://github.com/renovatebot/renovate).
<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4xMDEuMSIsInVwZGF0ZWRJblZlciI6IjQzLjEwMS4xIiwidGFyZ2V0QnJhbmNoIjoibWFpbiIsImxhYmVscyI6WyJyZW5vdmF0ZS9jb250YWluZXIiLCJ0eXBlL3BhdGNoIl19-->

Reviewed-on: https://git.erwanleboucher.dev/eleboucher/homelab/pulls/782
SysAdminDoc pushed a commit to SysAdminDoc/hermes-webui that referenced this pull request Jun 26, 2026
…lice doc + nesquena#3341 profile skill counts

nesquena#3327 fix(reasoning): normalize custom-provider model ids for fallback heuristics
Co-authored-by: Carry00 <Carry00@users.noreply.github.com>

nesquena#3334 docs(rfc): mark run-adapter Slice 4f shipped, define Slice 4g gate
Co-authored-by: Michaelyklam <Michaelyklam@users.noreply.github.com>

nesquena#3341 fix(profiles): show enabled vs compatible skill counts
Co-authored-by: b3nw <b3nw@users.noreply.github.com>
SysAdminDoc pushed a commit to SysAdminDoc/hermes-webui that referenced this pull request Jun 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants