Skip to content

feat: add LLM_CHEAP_MODEL for generic smart routing across all backends - #1081

Merged
ilblackdragon merged 2 commits into
nearai:mainfrom
smkrv:feat/generic-smart-routing-v2
Mar 16, 2026
Merged

ilblackdragon merged 2 commits into
nearai:mainfrom
smkrv:feat/generic-smart-routing-v2

Conversation

@smkrv

@smkrv smkrv commented Mar 12, 2026

Copy link
Copy Markdown
Contributor

Summary

Generalizes smart routing beyond NearAI — any LLM backend can now use a cheap/fast model for simple tasks via LLM_CHEAP_MODEL.

  • Generic LLM_CHEAP_MODEL env var works with any backend (NearAI, OpenAI, Anthropic, Ollama, OpenAI-compatible, Tinfoil)
  • Backward compatible: falls back to NEARAI_CHEAP_MODEL when backend is NearAI
  • Top-level SMART_ROUTING_CASCADE controls cascade behavior for all backends
  • Bedrock returns a clear error (smart routing not yet supported for native Converse API)
  • No .expect()/.unwrap() — all error paths use ok_or_else with LlmError

Resolution order

  1. LLM_CHEAP_MODEL (generic, any backend)
  2. NEARAI_CHEAP_MODEL (NearAI only, backward compat)

How it works

  • NearAI: clones NearAiConfig, swaps model
  • Registry providers (OpenAI, Anthropic, Ollama, etc.): clones RegistryProviderConfig, swaps model
  • Bedrock: returns LlmError::RequestFailed explaining it's not yet supported

Supersedes #379 — rebased and rewritten from scratch on current main to account for the provider registry refactor.

Changes

  • src/llm/config.rs — add cheap_model, smart_routing_cascade fields
  • src/config/llm.rs — parse LLM_CHEAP_MODEL, SMART_ROUTING_CASCADE env vars
  • src/llm/mod.rs — new create_cheap_provider_for_backend(), updated create_cheap_llm_provider() and build_provider_chain()
  • src/setup/wizard.rs — initialize new fields

Test plan

  • cargo clippy --all --benches --tests --examples --all-features — zero warnings
  • cargo fmt --check — clean
  • Existing cheap LLM tests updated to async + pass
  • New tests: generic override priority, NearAI fallback ignored for non-NearAI, Bedrock error path
  • Manual: set LLM_BACKEND=openai LLM_CHEAP_MODEL=gpt-4o-mini and verify smart routing log

Add generic cheap model support that works with any LLM backend, not just
NearAI. New env vars: LLM_CHEAP_MODEL (cheap model for any backend) and
SMART_ROUTING_CASCADE (top-level cascade flag).

Resolution order: LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL (backward compat).
Registry-based providers (OpenAI, Anthropic, Groq, etc.) clone their
RegistryProviderConfig with the cheap model swapped in. Bedrock returns
an explicit error (not yet supported). All error paths use ok_or_else
with proper LlmError variants -- no unwrap/expect in production code.
@github-actions github-actions Bot added scope: llm LLM integration scope: setup Onboarding / setup size: M 50-199 changed lines risk: high Safety, secrets, auth, or critical infrastructure labels Mar 12, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the smart routing capabilities by making the cheap model configuration generic across all LLM backends, rather than being limited to NearAI. It introduces new environment variables for global control over cheap model selection and smart routing cascade behavior, while maintaining backward compatibility for existing NearAI configurations. The changes streamline the process of setting up and utilizing cheap models for lightweight tasks, improving flexibility and future extensibility for various LLM providers.

Highlights

  • Generic LLM_CHEAP_MODEL: Introduced a new LLM_CHEAP_MODEL environment variable that allows any LLM backend (e.g., OpenAI, Anthropic, Ollama) to specify a cheap/fast model for simple tasks, generalizing smart routing beyond just NearAI.
  • Global SMART_ROUTING_CASCADE: Added a top-level SMART_ROUTING_CASCADE environment variable to control the cascade behavior for smart routing across all backends, defaulting to true.
  • Backward Compatibility and Priority: Ensured backward compatibility by falling back to NEARAI_CHEAP_MODEL when the backend is NearAI and LLM_CHEAP_MODEL is not set. The generic LLM_CHEAP_MODEL takes priority if both are configured.
  • Bedrock Support and Error Handling: Implemented specific error handling for Bedrock, returning a clear LlmError::RequestFailed when smart routing with a cheap model is attempted, as it is not yet supported for the native Converse API.
  • Refactored Cheap Provider Creation: Refactored the create_cheap_llm_provider function to be asynchronous and introduced a new create_cheap_provider_for_backend function to handle backend-specific logic for creating cheap LLM providers, improving modularity and extensibility.
Changelog
  • src/config/llm.rs
    • Added cheap_model and smart_routing_cascade fields to the default LlmConfig initialization.
    • Parsed LLM_CHEAP_MODEL and SMART_ROUTING_CASCADE environment variables into the LlmConfig.
  • src/llm/config.rs
    • Added cheap_model: Option<String> and smart_routing_cascade: bool fields to the LlmConfig struct definition with documentation.
  • src/llm/mod.rs
    • Updated create_cheap_llm_provider to be an asynchronous function and implemented the new cheap model resolution order (generic LLM_CHEAP_MODEL then NearAI-specific NEARAI_CHEAP_MODEL).
    • Introduced create_cheap_provider_for_backend to encapsulate backend-specific logic for creating cheap LLM providers, including handling NearAI, Bedrock, and other registry providers.
    • Modified build_provider_chain to utilize the new generic cheap model resolution and create_cheap_provider_for_backend.
    • Updated the cascade_enabled field in SmartRoutingConfig to use the new generic config.smart_routing_cascade.
    • Converted existing unit tests to #[tokio::test] async and added new tests for generic override priority, NearAI fallback for non-NearAI backends, and Bedrock error handling.
  • src/setup/wizard.rs
    • Initialized cheap_model: None and smart_routing_cascade: true in the default LlmConfig for the NearAI model fetch configuration.
Activity
  • The pull request was rebased and rewritten from scratch on the current main branch to integrate with the recent provider registry refactor, superseding a previous PR (feat: add LLM_CHEAP_MODEL for generic smart routing across all backends #379).
  • The author has performed comprehensive testing, including cargo clippy, cargo fmt, updating existing cheap LLM tests to async, and adding new tests for generic override priority, NearAI fallback, and Bedrock error paths.
  • Manual testing is planned to verify smart routing logs with LLM_BACKEND=openai LLM_CHEAP_MODEL=gpt-4o-mini.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request generalizes the smart routing feature by introducing a generic LLM_CHEAP_MODEL environment variable, which is a great enhancement for supporting various LLM backends. The changes are well-structured, and the backward compatibility for NearAI is maintained.

I have a couple of suggestions to improve maintainability and simplify the code:

  • Refactor duplicated logic for resolving the cheap model name into a helper method.
  • Remove unnecessary async from newly added functions to simplify the code.

Comment thread src/llm/mod.rs Outdated
Comment on lines 370 to 403
async fn create_cheap_provider_for_backend(
config: &LlmConfig,
session: Arc<SessionManager>,
cheap_model: &str,
) -> Result<Option<Arc<dyn LlmProvider>>, LlmError> {
if config.backend == "nearai" {
let mut cheap_config = config.nearai.clone();
cheap_config.model = cheap_model.to_string();
let provider =
create_llm_provider_with_config(&cheap_config, session, config.request_timeout_secs)?;
return Ok(Some(provider));
}

let mut cheap_config = config.nearai.clone();
cheap_config.model = cheap_model.clone();
if config.backend == "bedrock" {
return Err(LlmError::RequestFailed {
provider: "bedrock".to_string(),
reason: "Smart routing with cheap model is not supported for Bedrock yet".to_string(),
});
}

Ok(Some(Arc::new(NearAiChatProvider::new(
cheap_config,
session,
)?)))
// Registry-based provider: clone config and swap model
let reg_config = config.provider.as_ref().ok_or_else(|| LlmError::RequestFailed {
provider: config.backend.clone(),
reason: format!(
"Cannot create cheap provider for backend '{}': no registry provider config available",
config.backend
),
})?;

let mut cheap_reg_config = reg_config.clone();
cheap_reg_config.model = cheap_model.to_string();
let provider = create_registry_provider(&cheap_reg_config)?;
Ok(Some(provider))
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This function is marked as async but contains no .await calls, so it can be made synchronous. This would also allow create_cheap_llm_provider to be synchronous, simplifying the call sites in build_provider_chain by removing the need for .await. This would make the code slightly simpler and more efficient.

fn create_cheap_provider_for_backend(
    config: &LlmConfig,
    session: Arc<SessionManager>,
    cheap_model: &str,
) -> Result<Option<Arc<dyn LlmProvider>>, LlmError> {
    if config.backend == "nearai" {
        let mut cheap_config = config.nearai.clone();
        cheap_config.model = cheap_model.to_string();
        let provider =
            create_llm_provider_with_config(&cheap_config, session, config.request_timeout_secs)?;
        return Ok(Some(provider));
    }

    if config.backend == "bedrock" {
        return Err(LlmError::RequestFailed {
            provider: "bedrock".to_string(),
            reason: "Smart routing with cheap model is not supported for Bedrock yet".to_string(),
        });
    }

    // Registry-based provider: clone config and swap model
    let reg_config = config.provider.as_ref().ok_or_else(|| LlmError::RequestFailed {
        provider: config.backend.clone(),
        reason: format!(
            "Cannot create cheap provider for backend '{}': no registry provider config available",
            config.backend
        ),
    })?;

    let mut cheap_reg_config = reg_config.clone();
    cheap_reg_config.model = cheap_model.to_string();
    let provider = create_registry_provider(&cheap_reg_config)?;
    Ok(Some(provider))
}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in c70b49f. Removed async from both functions — neither had .await. Tests reverted to sync.

Comment thread src/llm/mod.rs Outdated
Comment on lines +451 to +457
let effective_cheap_model = config.cheap_model.as_deref().or_else(|| {
if config.backend == "nearai" {
config.nearai.cheap_model.as_deref()
} else {
None
}
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This logic for resolving the effective cheap model is duplicated from create_cheap_llm_provider (lines 349-355). To improve maintainability and follow the DRY principle, consider extracting this logic into a new method on LlmConfig. This aligns with the principle of minimizing boilerplate and maintaining consistency across configuration logic.

For example, you could add this to LlmConfig (e.g., in src/llm/config.rs):

impl LlmConfig {
    pub fn cheap_model_name(&self) -> Option<&str> {
        self.cheap_model.as_deref().or_else(|| {
            if self.backend == "nearai" {
                self.nearai.cheap_model.as_deref()
            } else {
                None
            }
        })
    }
}

Then you can simplify this block and the one in create_cheap_llm_provider to a single line.

    let effective_cheap_model = config.cheap_model_name();
References
  1. Use specialized helper functions when resolving configuration-related values to minimize boilerplate and maintain consistency.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in c70b49f. Extracted LlmConfig::cheap_model_name() in src/llm/config.rs — both create_cheap_llm_provider() and build_provider_chain() now use it. Added test_cheap_model_name_resolution() covering all 4 cases.

…heap_model_name()

- Remove async from create_cheap_provider_for_backend() and
  create_cheap_llm_provider() — neither contains .await calls
- Extract duplicated cheap model resolution logic into
  LlmConfig::cheap_model_name() helper method (DRY)
- Revert tests from tokio::test async back to sync #[test]
- Add test_cheap_model_name_resolution() unit test for the helper

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Clean generalization of cheap model support from NearAI-only to all registry-based backends. Resolution order (LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL) preserves backward compatibility. The cheap_model_name() helper on LlmConfig eliminates duplicated logic. Good test coverage. LGTM.

@ilblackdragon
ilblackdragon merged commit de214c2 into nearai:main Mar 16, 2026
22 checks passed
wehrmannit pushed a commit to wehrmannit/ironclaw that referenced this pull request Mar 20, 2026
The LLM_CHEAP_MODEL and SMART_ROUTING_CASCADE options from nearai#1081 were
only configurable via env vars. This adds them to the Settings struct
and web UI so users can configure smart routing from the browser.

Resolution order: env var > settings > default (None / true).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
…ds (nearai#1081)

* feat: add LLM_CHEAP_MODEL for generic smart routing across all backends

Add generic cheap model support that works with any LLM backend, not just
NearAI. New env vars: LLM_CHEAP_MODEL (cheap model for any backend) and
SMART_ROUTING_CASCADE (top-level cascade flag).

Resolution order: LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL (backward compat).
Registry-based providers (OpenAI, Anthropic, Groq, etc.) clone their
RegistryProviderConfig with the cheap model swapped in. Bedrock returns
an explicit error (not yet supported). All error paths use ok_or_else
with proper LlmError variants -- no unwrap/expect in production code.

* refactor: address Gemini review — remove unnecessary async, extract cheap_model_name()

- Remove async from create_cheap_provider_for_backend() and
  create_cheap_llm_provider() — neither contains .await calls
- Extract duplicated cheap model resolution logic into
  LlmConfig::cheap_model_name() helper method (DRY)
- Revert tests from tokio::test async back to sync #[test]
- Add test_cheap_model_name_resolution() unit test for the helper

---------

Co-authored-by: SMKRV <SMKRV@users.noreply.github.com>
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…ds (nearai#1081)

* feat: add LLM_CHEAP_MODEL for generic smart routing across all backends

Add generic cheap model support that works with any LLM backend, not just
NearAI. New env vars: LLM_CHEAP_MODEL (cheap model for any backend) and
SMART_ROUTING_CASCADE (top-level cascade flag).

Resolution order: LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL (backward compat).
Registry-based providers (OpenAI, Anthropic, Groq, etc.) clone their
RegistryProviderConfig with the cheap model swapped in. Bedrock returns
an explicit error (not yet supported). All error paths use ok_or_else
with proper LlmError variants -- no unwrap/expect in production code.

* refactor: address Gemini review — remove unnecessary async, extract cheap_model_name()

- Remove async from create_cheap_provider_for_backend() and
  create_cheap_llm_provider() — neither contains .await calls
- Extract duplicated cheap model resolution logic into
  LlmConfig::cheap_model_name() helper method (DRY)
- Revert tests from tokio::test async back to sync #[test]
- Add test_cheap_model_name_resolution() unit test for the helper

---------

Co-authored-by: SMKRV <SMKRV@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: regular 2-5 merged PRs risk: high Safety, secrets, auth, or critical infrastructure scope: llm LLM integration scope: setup Onboarding / setup size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants