feat: add LLM_CHEAP_MODEL for generic smart routing across all backends - #1081
Conversation
Add generic cheap model support that works with any LLM backend, not just NearAI. New env vars: LLM_CHEAP_MODEL (cheap model for any backend) and SMART_ROUTING_CASCADE (top-level cascade flag). Resolution order: LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL (backward compat). Registry-based providers (OpenAI, Anthropic, Groq, etc.) clone their RegistryProviderConfig with the cheap model swapped in. Bedrock returns an explicit error (not yet supported). All error paths use ok_or_else with proper LlmError variants -- no unwrap/expect in production code.
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly enhances the smart routing capabilities by making the cheap model configuration generic across all LLM backends, rather than being limited to NearAI. It introduces new environment variables for global control over cheap model selection and smart routing cascade behavior, while maintaining backward compatibility for existing NearAI configurations. The changes streamline the process of setting up and utilizing cheap models for lightweight tasks, improving flexibility and future extensibility for various LLM providers. Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request generalizes the smart routing feature by introducing a generic LLM_CHEAP_MODEL environment variable, which is a great enhancement for supporting various LLM backends. The changes are well-structured, and the backward compatibility for NearAI is maintained.
I have a couple of suggestions to improve maintainability and simplify the code:
- Refactor duplicated logic for resolving the cheap model name into a helper method.
- Remove unnecessary
asyncfrom newly added functions to simplify the code.
| async fn create_cheap_provider_for_backend( | ||
| config: &LlmConfig, | ||
| session: Arc<SessionManager>, | ||
| cheap_model: &str, | ||
| ) -> Result<Option<Arc<dyn LlmProvider>>, LlmError> { | ||
| if config.backend == "nearai" { | ||
| let mut cheap_config = config.nearai.clone(); | ||
| cheap_config.model = cheap_model.to_string(); | ||
| let provider = | ||
| create_llm_provider_with_config(&cheap_config, session, config.request_timeout_secs)?; | ||
| return Ok(Some(provider)); | ||
| } | ||
|
|
||
| let mut cheap_config = config.nearai.clone(); | ||
| cheap_config.model = cheap_model.clone(); | ||
| if config.backend == "bedrock" { | ||
| return Err(LlmError::RequestFailed { | ||
| provider: "bedrock".to_string(), | ||
| reason: "Smart routing with cheap model is not supported for Bedrock yet".to_string(), | ||
| }); | ||
| } | ||
|
|
||
| Ok(Some(Arc::new(NearAiChatProvider::new( | ||
| cheap_config, | ||
| session, | ||
| )?))) | ||
| // Registry-based provider: clone config and swap model | ||
| let reg_config = config.provider.as_ref().ok_or_else(|| LlmError::RequestFailed { | ||
| provider: config.backend.clone(), | ||
| reason: format!( | ||
| "Cannot create cheap provider for backend '{}': no registry provider config available", | ||
| config.backend | ||
| ), | ||
| })?; | ||
|
|
||
| let mut cheap_reg_config = reg_config.clone(); | ||
| cheap_reg_config.model = cheap_model.to_string(); | ||
| let provider = create_registry_provider(&cheap_reg_config)?; | ||
| Ok(Some(provider)) | ||
| } |
There was a problem hiding this comment.
This function is marked as async but contains no .await calls, so it can be made synchronous. This would also allow create_cheap_llm_provider to be synchronous, simplifying the call sites in build_provider_chain by removing the need for .await. This would make the code slightly simpler and more efficient.
fn create_cheap_provider_for_backend(
config: &LlmConfig,
session: Arc<SessionManager>,
cheap_model: &str,
) -> Result<Option<Arc<dyn LlmProvider>>, LlmError> {
if config.backend == "nearai" {
let mut cheap_config = config.nearai.clone();
cheap_config.model = cheap_model.to_string();
let provider =
create_llm_provider_with_config(&cheap_config, session, config.request_timeout_secs)?;
return Ok(Some(provider));
}
if config.backend == "bedrock" {
return Err(LlmError::RequestFailed {
provider: "bedrock".to_string(),
reason: "Smart routing with cheap model is not supported for Bedrock yet".to_string(),
});
}
// Registry-based provider: clone config and swap model
let reg_config = config.provider.as_ref().ok_or_else(|| LlmError::RequestFailed {
provider: config.backend.clone(),
reason: format!(
"Cannot create cheap provider for backend '{}': no registry provider config available",
config.backend
),
})?;
let mut cheap_reg_config = reg_config.clone();
cheap_reg_config.model = cheap_model.to_string();
let provider = create_registry_provider(&cheap_reg_config)?;
Ok(Some(provider))
}There was a problem hiding this comment.
Fixed in c70b49f. Removed async from both functions — neither had .await. Tests reverted to sync.
| let effective_cheap_model = config.cheap_model.as_deref().or_else(|| { | ||
| if config.backend == "nearai" { | ||
| config.nearai.cheap_model.as_deref() | ||
| } else { | ||
| None | ||
| } | ||
| }); |
There was a problem hiding this comment.
This logic for resolving the effective cheap model is duplicated from create_cheap_llm_provider (lines 349-355). To improve maintainability and follow the DRY principle, consider extracting this logic into a new method on LlmConfig. This aligns with the principle of minimizing boilerplate and maintaining consistency across configuration logic.
For example, you could add this to LlmConfig (e.g., in src/llm/config.rs):
impl LlmConfig {
pub fn cheap_model_name(&self) -> Option<&str> {
self.cheap_model.as_deref().or_else(|| {
if self.backend == "nearai" {
self.nearai.cheap_model.as_deref()
} else {
None
}
})
}
}Then you can simplify this block and the one in create_cheap_llm_provider to a single line.
let effective_cheap_model = config.cheap_model_name();References
- Use specialized helper functions when resolving configuration-related values to minimize boilerplate and maintain consistency.
There was a problem hiding this comment.
Fixed in c70b49f. Extracted LlmConfig::cheap_model_name() in src/llm/config.rs — both create_cheap_llm_provider() and build_provider_chain() now use it. Added test_cheap_model_name_resolution() covering all 4 cases.
…heap_model_name() - Remove async from create_cheap_provider_for_backend() and create_cheap_llm_provider() — neither contains .await calls - Extract duplicated cheap model resolution logic into LlmConfig::cheap_model_name() helper method (DRY) - Revert tests from tokio::test async back to sync #[test] - Add test_cheap_model_name_resolution() unit test for the helper
zmanian
left a comment
There was a problem hiding this comment.
Clean generalization of cheap model support from NearAI-only to all registry-based backends. Resolution order (LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL) preserves backward compatibility. The cheap_model_name() helper on LlmConfig eliminates duplicated logic. Good test coverage. LGTM.
The LLM_CHEAP_MODEL and SMART_ROUTING_CASCADE options from nearai#1081 were only configurable via env vars. This adds them to the Settings struct and web UI so users can configure smart routing from the browser. Resolution order: env var > settings > default (None / true). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ds (nearai#1081) * feat: add LLM_CHEAP_MODEL for generic smart routing across all backends Add generic cheap model support that works with any LLM backend, not just NearAI. New env vars: LLM_CHEAP_MODEL (cheap model for any backend) and SMART_ROUTING_CASCADE (top-level cascade flag). Resolution order: LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL (backward compat). Registry-based providers (OpenAI, Anthropic, Groq, etc.) clone their RegistryProviderConfig with the cheap model swapped in. Bedrock returns an explicit error (not yet supported). All error paths use ok_or_else with proper LlmError variants -- no unwrap/expect in production code. * refactor: address Gemini review — remove unnecessary async, extract cheap_model_name() - Remove async from create_cheap_provider_for_backend() and create_cheap_llm_provider() — neither contains .await calls - Extract duplicated cheap model resolution logic into LlmConfig::cheap_model_name() helper method (DRY) - Revert tests from tokio::test async back to sync #[test] - Add test_cheap_model_name_resolution() unit test for the helper --------- Co-authored-by: SMKRV <SMKRV@users.noreply.github.com>
…ds (nearai#1081) * feat: add LLM_CHEAP_MODEL for generic smart routing across all backends Add generic cheap model support that works with any LLM backend, not just NearAI. New env vars: LLM_CHEAP_MODEL (cheap model for any backend) and SMART_ROUTING_CASCADE (top-level cascade flag). Resolution order: LLM_CHEAP_MODEL > NEARAI_CHEAP_MODEL (backward compat). Registry-based providers (OpenAI, Anthropic, Groq, etc.) clone their RegistryProviderConfig with the cheap model swapped in. Bedrock returns an explicit error (not yet supported). All error paths use ok_or_else with proper LlmError variants -- no unwrap/expect in production code. * refactor: address Gemini review — remove unnecessary async, extract cheap_model_name() - Remove async from create_cheap_provider_for_backend() and create_cheap_llm_provider() — neither contains .await calls - Extract duplicated cheap model resolution logic into LlmConfig::cheap_model_name() helper method (DRY) - Revert tests from tokio::test async back to sync #[test] - Add test_cheap_model_name_resolution() unit test for the helper --------- Co-authored-by: SMKRV <SMKRV@users.noreply.github.com>
Summary
Generalizes smart routing beyond NearAI — any LLM backend can now use a cheap/fast model for simple tasks via
LLM_CHEAP_MODEL.LLM_CHEAP_MODELenv var works with any backend (NearAI, OpenAI, Anthropic, Ollama, OpenAI-compatible, Tinfoil)NEARAI_CHEAP_MODELwhen backend is NearAISMART_ROUTING_CASCADEcontrols cascade behavior for all backends.expect()/.unwrap()— all error paths useok_or_elsewithLlmErrorResolution order
LLM_CHEAP_MODEL(generic, any backend)NEARAI_CHEAP_MODEL(NearAI only, backward compat)How it works
NearAiConfig, swaps modelRegistryProviderConfig, swaps modelLlmError::RequestFailedexplaining it's not yet supportedSupersedes #379 — rebased and rewritten from scratch on current
mainto account for the provider registry refactor.Changes
src/llm/config.rs— addcheap_model,smart_routing_cascadefieldssrc/config/llm.rs— parseLLM_CHEAP_MODEL,SMART_ROUTING_CASCADEenv varssrc/llm/mod.rs— newcreate_cheap_provider_for_backend(), updatedcreate_cheap_llm_provider()andbuild_provider_chain()src/setup/wizard.rs— initialize new fieldsTest plan
cargo clippy --all --benches --tests --examples --all-features— zero warningscargo fmt --check— cleanLLM_BACKEND=openai LLM_CHEAP_MODEL=gpt-4o-miniand verify smart routing log