Skip to content

feat(web): expose cheap_model and smart_routing_cascade in settings UI - #1491

Closed
wehrmannit wants to merge 2 commits into
nearai:mainfrom
wehrmannit:feat/web-cheap-model-settings
Closed

wehrmannit wants to merge 2 commits into
nearai:mainfrom
wehrmannit:feat/web-cheap-model-settings

Conversation

@wehrmannit

@wehrmannit wehrmannit commented Mar 20, 2026 •

Copy link
Copy Markdown

Summary

  • Adds cheap_model and smart_routing_cascade fields to the Settings struct so they persist via settings.json / DB
  • Exposes both in the web UI settings page (Inference tab) with i18n support (en + zh-CN)
  • Updates LlmConfig::resolve() to fall back to settings when env vars aren't set (env var > settings > default)
  • Marks both keys as restart-required in the UI

Follows up on #1081 which added LLM_CHEAP_MODEL and SMART_ROUTING_CASCADE env vars but left them unconfigurable from the browser.

Test plan

  • cargo clippy --all --benches --tests --examples --all-features — zero warnings
  • cargo build --release — compiles clean
  • Open web UI Settings > Inference, verify "Cheap Model" text field and "Smart Routing Cascade" toggle appear
  • Set a cheap model via the UI, restart, confirm it takes effect in smart routing logs
  • Verify env var still overrides the UI setting

🤖 Generated with Claude Code

The LLM_CHEAP_MODEL and SMART_ROUTING_CASCADE options from nearai#1081 were
only configurable via env vars. This adds them to the Settings struct
and web UI so users can configure smart routing from the browser.

Resolution order: env var > settings > default (None / true).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added scope: channel/web Web gateway channel scope: config Configuration size: S 10-49 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: new First-time contributor labels Mar 20, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enhances the application's LLM configuration capabilities by integrating previously environment-variable-only settings, cheap_model and smart_routing_cascade, into the persistent settings system and exposing them through the web user interface. This allows administrators to manage these crucial smart routing parameters directly from the UI, improving configurability and user experience, while maintaining environment variable precedence.

Highlights

  • Persistent Settings: Introduced cheap_model and smart_routing_cascade fields into the Settings struct for persistent storage via settings.json / DB.
  • Web UI Exposure: Exposed new configuration options for "Cheap Model" (text field) and "Smart Routing Cascade" (toggle) in the web UI's Inference settings, including internationalization for English and Simplified Chinese.
  • Configuration Resolution Logic: Implemented a configuration resolution logic in LlmConfig::resolve() that prioritizes environment variables, then settings from the UI/DB, and finally falls back to default values for cheap_model and smart_routing_cascade.
  • Restart Requirement: Marked the new UI settings for cheap_model and smart_routing_cascade as requiring a restart to take effect.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces new configuration options for LLM smart routing: cheap_model and smart_routing_cascade. The cheap_model allows specifying a model for lightweight tasks, while smart_routing_cascade enables a retry mechanism with the primary model if the cheap model's response is uncertain. These settings are added to the application's configuration structure, integrated into the LLM configuration loading logic (prioritizing environment variables over settings), and exposed in the web UI with corresponding internationalization strings. Changes to these settings will require an application restart.

@serrrfirat serrrfirat left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Paranoid Architect Review - All Findings

🔴 High Severity Issues

1. Type Safety - No validation for cheap_model format

Location: src/config/llm.rs:216-223

The resolve() method reads cheap_model from Settings but there's no validation that it's a valid provider/model combination. Invalid values will cause runtime errors during LLM initialization.

Recommendation: Add validation in Settings deserialization or in resolve() to verify the cheap_model string format matches expected provider:model pattern (e.g., "openai:gpt-4o-mini"). Consider adding a helper method to validate model strings against available providers.

2. Database Compatibility - Missing dual-backend verification

Location: src/settings.rs:87-93

New Optional fields added to Settings struct but no database migration or dual-backend verification. Per CLAUDE.md: "All new persistence features must support both backends." Settings are persisted via db.set_settings_map() - need to verify both PostgreSQL and libSQL handle these new fields.

Recommendation: Verify that both PostgreSQL and libSQL backends correctly serialize/deserialize the new Optional and Optional fields. Add integration tests with #[cfg(feature = "integration")] to validate Settings persistence across both backends.


🟡 Medium Severity Issues

3. User Experience - No client-side validation

Location: src/channels/web/static/app.js:4708-4709

New settings marked as RESTART_REQUIRED but no validation of cheap_model format in the UI. Users can enter arbitrary strings that will cause silent failures or cryptic errors on restart.

Recommendation: Add client-side validation to ensure cheap_model matches expected format (provider:model). Show available models in a dropdown or add format hint text. Validate on blur/change before allowing save.

4. Configuration Priority - Asymmetric defaults

Location: src/config/llm.rs:216-223

The resolution chain is env var → settings → default but there's asymmetry: cheap_model has no default while smart_routing_cascade defaults to true. This creates inconsistent behavior - smart routing might be enabled but fail if no cheap model is configured.

Recommendation: Add defensive logic: if smart_routing_cascade resolves to true but cheap_model is None, either (1) emit a warning and disable cascade, or (2) use the primary model as cheap model. Document the expected behavior when one is set without the other.

5. Documentation - Missing interaction explanation

Location: src/settings.rs:87-92, src/config/llm.rs:216-223

Comments explain what each field does but don't explain the interaction between them. What happens if cascade is true but cheap_model is None? What's the expected format for cheap_model?

Recommendation: Add a module-level doc comment in settings.rs explaining the smart routing feature, the relationship between these two fields, expected cheap_model format (provider:model), and fallback behavior.


🟢 Low Severity Issues

6. Code Quality - Growing resolve() method

Location: src/config/llm.rs:216-223

The resolve() method is growing large with more conditional logic. Adding two more optional fields continues the pattern of manual env-then-settings resolution.

Recommendation: Consider extracting a helper method like resolve_optional<T>(env_key: &str, setting: Option<T>, default: T) -> T to reduce duplication. This PR is fine as-is but sets precedent for future refactoring.

7. Internationalization - Brief translation strings

Location: src/channels/web/static/i18n/zh-CN.js, en.js

Added translations for the new settings but translations are brief. The English description for cheap_model is just "Cheap model for smart routing" which doesn't explain the format or give examples.

Recommendation: Enhance translation strings to include format hints, e.g., "Cheap model for smart routing (format: provider:model, e.g., openai:gpt-4o-mini)". This helps international users understand expected input.


Summary: The PR correctly exposes two LLM configuration options to the Settings UI, but lacks validation for the cheap_model format and needs dual-backend database verification per project requirements. The most critical issues are: (1) no validation that cheap_model is a valid provider:model string, and (2) no verification that PostgreSQL and libSQL both handle the new Optional fields correctly.

@serrrfirat

Copy link
Copy Markdown
Collaborator

I did a paranoid pass against current staging and pushed a hardened implementation here:

  • branch: serrrfirat:firat/pr-1491-paranoid-fixes
  • commit: f8d78093 (fix(web): harden cheap-model settings rollout)

Why I think the original shape needs tightening in current staging:

  1. cheap_model and smart_routing_cascade affect the shared LLM provider chain, so they need to participate in the existing hot-reload path.
  2. They should be admin-scoped, not per-user settings.
  3. owner/admin checks in the settings surface should use user.is_admin(), not raw role == "admin".
  4. smart_routing_cascade needs an inherit/unset UI state to preserve env > settings > default, so a plain boolean toggle is misleading.
  5. The frontend has moved from the old monolithic src/channels/web/static/app.js path to crates/ironclaw_gateway/static/js/surfaces/settings.js, so the patch needs to land there.

What’s in the pushed branch:

  • adds Settings.cheap_model and Settings.smart_routing_cascade
  • resolves them as env > settings > default in LlmConfig::resolve()
  • adds both keys to admin-only gating and LLM hot-reload trigger handling
  • fixes settings admin checks to use user.is_admin()
  • makes /api/settings/export support ?scope=admin
  • refreshes gateway status snapshot with cheap_model and smart_routing_cascade
  • updates the current web settings UI to:
    • load admin-scope inference settings separately
    • save them with ?scope=admin
    • show a tri-state-ish cascade control (inherit / enabled / disabled via select)
  • adds i18n strings (en, zh-CN)
  • adds targeted tests for config precedence and admin-scope export access

Targeted validation I ran locally:

  • cargo fmt --all
  • cargo test --lib cheap_model_uses_settings_when_env_unset
  • cargo test --lib smart_routing_cascade_uses_settings_when_env_unset
  • cargo test --lib test_settings_export_owner_can_read_admin_scope
  • cargo test --lib settings_set_handler_triggers_llm_provider_hot_reload
  • node --check crates/ironclaw_gateway/static/js/surfaces/settings.js
  • node --check crates/ironclaw_gateway/static/i18n/en.js
  • node --check crates/ironclaw_gateway/static/i18n/zh-CN.js

If useful, I can also open a follow-up PR from my branch with just these fixes layered on top of #1491.

@henrypark133
henrypark133 changed the base branch from staging to main May 1, 2026 06:19
@serrrfirat

Copy link
Copy Markdown
Collaborator

Thank you for making smart-routing controls available in the browser. We are closing this PR because it modifies the retired v1 settings and WebUI implementation.

Reborn retains the underlying cheap_model and smart_routing_cascade configuration in crates/ironclaw_llm, but user-facing configuration now needs to flow through the Reborn LLM administration facade and WebUI v2 rather than src/channels/web/static and the v1 Settings type. Forward-porting this patch would create a second configuration source.

We would be glad to have you port this experience to the Reborn administration API and WebUI v2, including persistence, precedence, and caller-level tests. Thank you for addressing an important usability gap.

@serrrfirat serrrfirat closed this Jul 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: new First-time contributor risk: medium Business logic, config, or moderate-risk modules scope: channel/web Web gateway channel scope: config Configuration size: S 10-49 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants