feat(sdk): Add advanced orchestration of model and providers BZ-43839 - #146
Conversation
|
Important Review skippedAuto incremental reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the You can disable this status message by setting the WalkthroughAdds Advanced Orchestration docs and introduces a data-driven classification and routing system: configuration constants, prompt analysis utilities, a binary classifier, and a model router. Integrates orchestration into NeuroLink with a new enableOrchestration flag and expands MCP server/tooling APIs, logging, diagnostics, and fallbacks. Changes
Sequence Diagram(s)sequenceDiagram
autonumber
actor User
participant NL as NeuroLink
participant PE as Precedence Engine
participant BTC as BinaryTaskClassifier
participant MR as ModelRouter
participant Provider as Provider API
User->>NL: generate(prompt, options)
NL->>PE: Resolve explicit provider/model?
alt Explicit specified
PE-->>NL: Use specified route
else Orchestration enabled
NL->>BTC: classify(prompt)
BTC-->>NL: {type, confidence, reasoning}
NL->>MR: route(prompt, constraints)
MR-->>NL: {provider, model, confidence}
else Orchestration disabled
NL-->>NL: Use original options
end
NL->>Provider: request(provider, model, prompt)
alt Success
Provider-->>NL: response
NL-->>User: response
else Error
NL->>MR: getFallbackRoute(...)
MR-->>NL: fallback route
NL->>Provider: request(fallback)
Provider-->>NL: response or error
NL-->>User: response or propagated error
end
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
✨ Finishing Touches🧪 Generate unit tests
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
e1de1d4 to
3fa7b7b
Compare
There was a problem hiding this comment.
Actionable comments posted: 11
🧹 Nitpick comments (16)
docs/ADVANCED-ORCHESTRATION.md (2)
51-56: Avoid hard-coding dated model IDs in docs; prefer config-driven placeholders.The string "claude-sonnet-4@20250514" will age quickly and may mislead users. Suggest removing the date suffix or referencing a config variable (e.g., REASONING_PRIMARY_MODEL) in examples and the flow diagram/debug lines.
-// → Uses vertex/claude-sonnet-4@20250514 +// → Uses vertex/claude-sonnet-4 (exact version from config) ... -Model: gemini-2.5-flash | claude-sonnet-4@20250514 +Model: gemini-2.5-flash | claude-sonnet-4 ... -// [DEBUG] Orchestration applied: reasoning -> vertex/claude-sonnet-4@20250514 +// [DEBUG] Orchestration applied: reasoning -> vertex/claude-sonnet-4Also applies to: 187-190, 227-230
201-207: Qualification needed for perf claims (<10ms classify, <5ms route).Unless these are measured across CI or a benchmark suite, add “target”/“typical” language and a note about environment variance, or link to a benchmark result.
Also applies to: 367-368
src/lib/config/taskClassificationConfig.ts (4)
16-29: Broaden “fast” regexes to handle multi-word subjects and punctuation.Current patterns like
^what is\s+\w+and^tell me about\s+\w+miss multi-word entities ("what is neural network pruning?"). Loosen them to accept phrases.- /^what is\s+\w+\??$/i, + /^what\s+is\s+.+\??$/i, - /^tell me about\s+\w+$/i, + /^tell\s+me\s+about\s+.+$/i, - /^what does\s+\w+\s+mean/i, + /^what\s+does\s+.+?\s+mean\??/i, - /^translate\s+["'].*["']\s+to\s+\w+/i, + /^translate\s+["'].+["']\s+to\s+\w+/i, - /^how do you say\s+/i, + /^how\s+do\s+you\s+say\s+.+\??/i,Also applies to: 36-38
83-103: Augment fast keywords for parity with docs.Docs cite “time, weather, translate.” Consider adding these to FAST_KEYWORDS for better recall when keyword-based scoring triggers.
"count", + "time", + "date", + "weather", + "translate",
151-158: Confidence clamping may overstate borderline cases.MIN_CONFIDENCE=0.6 forces 0.51 ratios up to 0.6. If that’s intentional, fine; otherwise consider lowering to 0.5 or removing the floor to reflect true uncertainty.
163-166: Capture “what’s” contractions in SIMPLE_DEFINITION.Add a contraction-friendly alternative for “what’s/whats …”.
export const DOMAIN_PATTERNS = { TECHNICAL: /\b(code|programming|development|software)\b/i, - SIMPLE_DEFINITION: /\b(definition|meaning|what is)\b/i, + SIMPLE_DEFINITION: /\b(definition|meaning|what\s+is|what['’]?s)\b/i, } as const;src/lib/utils/modelRouter.ts (3)
34-86: Externalize model catalog (IDs, latency, cost) and add region-aware availability.Hard-coded model IDs, costs, and latencies will drift. Move MODEL_CONFIGS to a config file (e.g., src/lib/config/modelRoutingConfig.ts), read region (GOOGLE_CLOUD_LOCATION), and pick only available models. Provide a soft fallback across providers when Claude isn’t in-region.
262-282: Async function without awaits.validateRoute is async but fully synchronous. Either remove async or add an actual availability check when feasible.
184-223: Fallback strategy could consider cross-provider options.“auto” flips task type; sometimes staying in-type with another provider is better (e.g., fast→fast on OpenAI if Vertex issue). Consider a provider-aware fallback matrix.
src/lib/utils/taskClassificationUtils.ts (4)
50-56: Accumulate all fast-pattern matches (don’t early-return).Returning on first hit loses signal strength and reasoning context. Count matches and scale score; push count in reasons.
export function checkFastPatterns( normalizedPrompt: string, reasons: string[], ): number { - for (const pattern of FAST_PATTERNS) { - if (pattern.test(normalizedPrompt)) { - reasons.push("fast pattern match"); - return SCORING_WEIGHTS.PATTERN_MATCH_SCORE; - } - } - return 0; + let matches = 0; + for (const pattern of FAST_PATTERNS) { + if (pattern.test(normalizedPrompt)) matches++; + } + if (matches > 0) { + reasons.push(`${matches} fast pattern match${matches > 1 ? "es" : ""}`); + } + return matches * SCORING_WEIGHTS.PATTERN_MATCH_SCORE; }
66-73: Mirror reasoning-pattern logic to count all matches.Same rationale as fast patterns; improves classifier granularity.
export function checkReasoningPatterns( normalizedPrompt: string, reasons: string[], ): number { - for (const pattern of REASONING_PATTERNS) { - if (pattern.test(normalizedPrompt)) { - reasons.push("reasoning pattern match"); - return SCORING_WEIGHTS.PATTERN_MATCH_SCORE; - } - } - return 0; + let matches = 0; + for (const pattern of REASONING_PATTERNS) { + if (pattern.test(normalizedPrompt)) matches++; + } + if (matches > 0) { + reasons.push(`${matches} reasoning pattern match${matches > 1 ? "es" : ""}`); + } + return matches * SCORING_WEIGHTS.PATTERN_MATCH_SCORE; }
90-93: Use const for non-reassigned locals (lint warnings).Removes warnings and signals immutability.
- let fastScore = fastKeywordMatches * SCORING_WEIGHTS.KEYWORD_MATCH_SCORE; - let reasoningScore = + const fastScore = fastKeywordMatches * SCORING_WEIGHTS.KEYWORD_MATCH_SCORE; + const reasoningScore = reasoningKeywordMatches * SCORING_WEIGHTS.KEYWORD_MATCH_SCORE;
197-233: Add unit samples to guard regressions.Please add quick tests for short greeting, multi-question “why + how”, and “what is ” under SIMPLE_DEFINITION_LENGTH to confirm scoring and tie-break behavior post-change.
src/lib/neurolink.ts (3)
108-110: Avoid double classification; drop BinaryTaskClassifier import if unused after fix.If we remove the extra classify() call in applyOrchestration (see below), this import can go.
207-226: Document regional requirement for Claude-on-Vertex.Add note that GOOGLE_CLOUD_LOCATION=us-east5 is required for Claude via Vertex AI; include quick troubleshooting tip.
241-243: Also log orchestration enablement at constructor start.Helps correlate routing behavior in logs.
this.logConstructorStart( constructorId, constructorStartTime, constructorHrTimeStart, - config, + { ...(config || {}), enableOrchestration: this.enableOrchestration } as any, );
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
💡 Knowledge Base configuration:
- MCP integration is disabled by default for public repositories
- Jira integration is disabled by default for public repositories
- Linear integration is disabled by default for public repositories
You can enable these sources in your CodeRabbit configuration.
📒 Files selected for processing (6)
docs/ADVANCED-ORCHESTRATION.md(1 hunks)src/lib/config/taskClassificationConfig.ts(1 hunks)src/lib/neurolink.ts(7 hunks)src/lib/utils/modelRouter.ts(1 hunks)src/lib/utils/taskClassificationUtils.ts(1 hunks)src/lib/utils/taskClassifier.ts(1 hunks)
🧰 Additional context used
🧬 Code graph analysis (4)
src/lib/utils/taskClassifier.ts (3)
src/lib/utils/taskClassificationUtils.ts (4)
ClassificationScores(16-20)analyzePrompt(197-233)determineTaskType(186-191)calculateConfidence(166-181)src/lib/config/taskClassificationConfig.ts (1)
CLASSIFICATION_THRESHOLDS(151-158)src/lib/utils/logger.ts (1)
logger(341-380)
src/lib/utils/modelRouter.ts (2)
src/lib/utils/taskClassifier.ts (3)
TaskType(15-15)TaskClassification(17-21)BinaryTaskClassifier(27-118)src/lib/utils/logger.ts (2)
logger(341-380)error(223-225)
src/lib/utils/taskClassificationUtils.ts (1)
src/lib/config/taskClassificationConfig.ts (7)
CLASSIFICATION_THRESHOLDS(151-158)SCORING_WEIGHTS(137-146)FAST_PATTERNS(9-38)REASONING_PATTERNS(43-78)FAST_KEYWORDS(83-103)REASONING_KEYWORDS(108-132)DOMAIN_PATTERNS(163-166)
src/lib/neurolink.ts (4)
src/lib/types/generateTypes.ts (1)
GenerateOptions(14-61)src/lib/utils/modelRouter.ts (2)
route(95-179)ModelRouter(91-356)src/lib/utils/logger.ts (2)
logger(341-380)error(223-225)src/lib/utils/taskClassifier.ts (1)
BinaryTaskClassifier(27-118)
🪛 LanguageTool
docs/ADVANCED-ORCHESTRATION.md
[grammar] ~9-~9: There might be a mistake here.
Context: ...tures ### 🧠 Binary Task Classification - Fast Tasks: Simple queries, calculatio...
(QB_NEW_EN)
[grammar] ~11-~11: There might be a mistake here.
Context: ...s → Routed to Vertex AI Gemini 2.5 Flash - Reasoning Tasks: Complex analysis, phi...
(QB_NEW_EN)
[grammar] ~16-~16: There might be a mistake here.
Context: ...r and model selection based on task type - Optimizes for response speed vs. reasoni...
(QB_NEW_EN)
[grammar] ~17-~17: There might be a mistake here.
Context: ... response speed vs. reasoning capability - Built-in confidence scoring for classifi...
(QB_NEW_EN)
[grammar] ~20-~20: There might be a mistake here.
Context: ...on accuracy ### 🎯 Precedence Hierarchy 1. User-specified provider/model (highest...
(QB_NEW_EN)
[grammar] ~22-~22: There might be a mistake here.
Context: ...fied provider/model** (highest priority) 2. Orchestration routing (when no provide...
(QB_NEW_EN)
[grammar] ~23-~23: There might be a mistake here.
Context: ...n routing** (when no provider specified) 3. Auto provider selection (fallback) 4. ...
(QB_NEW_EN)
[grammar] ~24-~24: There might be a mistake here.
Context: .... Auto provider selection (fallback) 4. Graceful error handling ### 🔄 Zero B...
(QB_NEW_EN)
[grammar] ~27-~27: There might be a mistake here.
Context: ...handling** ### 🔄 Zero Breaking Changes - Completely optional feature (disabled by...
(QB_NEW_EN)
[grammar] ~29-~29: There might be a mistake here.
Context: ...y optional feature (disabled by default) - Existing functionality preserved - Backw...
(QB_NEW_EN)
[grammar] ~30-~30: There might be a mistake here.
Context: ...ault) - Existing functionality preserved - Backward compatible with all existing co...
(QB_NEW_EN)
[grammar] ~98-~98: There might be a mistake here.
Context: ...) - Short prompts (< 50 characters) - Keywords: quick, fast, simple, what, t...
(QB_NEW_EN)
[grammar] ~99-~99: There might be a mistake here.
Context: ...hat, time, weather, calculate, translate - Patterns: Questions, calculations, gre...
(QB_NEW_EN)
[grammar] ~100-~100: There might be a mistake here.
Context: ...calculations, greetings, simple requests - Examples: - "What's 2+2?" - "Curre...
(QB_NEW_EN)
[grammar] ~101-~101: There might be a mistake here.
Context: ...eetings, simple requests - Examples: - "What's 2+2?" - "Current time?" - "Q...
(QB_NEW_EN)
[grammar] ~102-~102: There might be a mistake here.
Context: ...quests - Examples: - "What's 2+2?" - "Current time?" - "Quick weather updat...
(QB_NEW_EN)
[grammar] ~109-~109: There might be a mistake here.
Context: ...x prompts** (detailed analysis requests) - Keywords: analyze, explain, compare, d...
(QB_NEW_EN)
[grammar] ~110-~110: There might be a mistake here.
Context: ...ategy, implications, philosophy, complex - Patterns: Analysis requests, philosoph...
(QB_NEW_EN)
[grammar] ~111-~111: There might be a mistake here.
Context: ...sophical questions, strategy development - Examples: - "Analyze the ethical imp...
(QB_NEW_EN)
[grammar] ~112-~112: There might be a mistake here.
Context: ...ns, strategy development - Examples: - "Analyze the ethical implications of AI ...
(QB_NEW_EN)
[grammar] ~168-~168: There might be a mistake here.
Context: ...ation**: Orchestration logic integrated into main generation flow 4. **Precedence En...
(QB_NEW_EN)
[grammar] ~194-~194: There might be a mistake here.
Context: ...*: Falls back to auto provider selection - Provider Unavailable: Uses next best a...
(QB_NEW_EN)
[grammar] ~195-~195: There might be a mistake here.
Context: ...ble**: Uses next best available provider - Classification Errors: Defaults to fas...
(QB_NEW_EN)
[grammar] ~196-~196: There might be a mistake here.
Context: ... Errors**: Defaults to fast task routing - Network Issues: Standard NeuroLink ret...
(QB_NEW_EN)
[grammar] ~253-~253: There might be a mistake here.
Context: ...kloads (both simple and complex queries) - Cost optimization important - Response t...
(QB_NEW_EN)
[grammar] ~254-~254: There might be a mistake here.
Context: ...x queries) - Cost optimization important - Response time optimization for simple qu...
(QB_NEW_EN)
[grammar] ~255-~255: There might be a mistake here.
Context: ...nse time optimization for simple queries - Large-scale applications with varied req...
(QB_NEW_EN)
[grammar] ~260-~260: There might be a mistake here.
Context: ...applications (all fast or all reasoning) - When you need consistent provider behavi...
(QB_NEW_EN)
[grammar] ~261-~261: There might be a mistake here.
Context: ...en you need consistent provider behavior - Testing/development with specific models...
(QB_NEW_EN)
[grammar] ~262-~262: There might be a mistake here.
Context: ...Testing/development with specific models - Applications requiring strict provider c...
(QB_NEW_EN)
[grammar] ~431-~431: There might be a mistake here.
Context: ...implementation of Advanced Orchestration - Binary task classification - Intellige...
(QB_NEW_EN)
[grammar] ~432-~432: There might be a mistake here.
Context: ...estration - Binary task classification - Intelligent model routing - Zero break...
(QB_NEW_EN)
[grammar] ~433-~433: There might be a mistake here.
Context: ...sification - Intelligent model routing - Zero breaking changes - Comprehensive ...
(QB_NEW_EN)
[grammar] ~434-~434: There might be a mistake here.
Context: ... model routing - Zero breaking changes - Comprehensive testing and validation ##...
(QB_NEW_EN)
🪛 GitHub Check: test (18)
src/lib/utils/taskClassificationUtils.ts
[warning] 91-91:
'reasoningScore' is never reassigned. Use 'const' instead
[warning] 90-90:
'fastScore' is never reassigned. Use 'const' instead
🪛 GitHub Check: 🛡️ Code Quality & Security Gate
src/lib/utils/taskClassificationUtils.ts
[warning] 91-91:
'reasoningScore' is never reassigned. Use 'const' instead
[warning] 90-90:
'fastScore' is never reassigned. Use 'const' instead
🪛 GitHub Check: test (20)
src/lib/utils/taskClassificationUtils.ts
[warning] 91-91:
'reasoningScore' is never reassigned. Use 'const' instead
[warning] 90-90:
'fastScore' is never reassigned. Use 'const' instead
🔇 Additional comments (6)
docs/ADVANCED-ORCHESTRATION.md (2)
279-284: enableAnalytics and getEventEmitter confirmed present and exported; docs accurate.
431-436: Versions are up-to-date
package.json is at v7.31.0 and CHANGELOG.md includes the 7.31.0 entry, matching the docs.src/lib/utils/taskClassifier.ts (2)
31-52: Classification control flow looks solid.Defaulting to “fast” only when both scores are zero is reasonable; otherwise relying on determineTaskType/calculateConfidence is clear and testable.
6-13: ESM .js imports are supported — tsconfig.json is configured with"module": "NodeNext"and"moduleResolution": "NodeNext", so the.jsimports in.tsfiles will resolve correctly.src/lib/neurolink.ts (2)
196-199: Good: explicit orchestration flag with sane default.
235-236: Constructor type expanded correctly.
c023857 to
e9f4c9d
Compare
494b9a2 to
993d878
Compare
993d878 to
a269482
Compare
|
🎉 This PR is included in version 7.37.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
Pull Request
Description
Complete implementation of Advanced Model Orchestration POC system with binary task classification. This POC introduces intelligent routing between fast and reasoning models, featuring Claude 4 and Gemini 2.5 flash integration via Vertex AI. The system implements a comprehensive task classification engine that analyzes prompts and routes them to optimal models for significant cost optimization and performance improvements.
Type of Change
Changes Made
Implemented complete task classification system with binary routing logic
Created src/lib/config/taskClassificationConfig.ts - centralized patterns, keywords, scoring weights, and classification thresholds
Created src/lib/utils/taskClassificationUtils.ts - comprehensive utility functions for prompt analysis and classification
Created enhanced src/lib/utils/taskClassifier.ts - core binary classification engine with advanced scoring
Implemented src/lib/utils/modelRouter.ts - intelligent model selection and routing system
Added fast task routing to Gemini 2.5 Flash via Vertex AI for cost optimization
Added reasoning task routing to Claude Sonnet 4 via Vertex AI for complex analysis
Implemented test-orchestration-poc.js - comprehensive POC validation and testing
Created docs/ADVANCED-ORCHESTRATION.md - complete documentation for the orchestration system
Added regional configuration support for Claude model availability
Implemented cost optimization analytics and performance tracking
AI Provider Impact
Component Impact
Testing
Test Environment
Performance Impact
Breaking Changes
None. This is a new POC feature with full backward compatibility. Requires new environment variable GOOGLE_CLOUD_LOCATION=us-east5 for Claude model access.
Checklist
Additional Notes
This POC demonstrates advanced model orchestration capabilities with 60-90% cost optimization potential. The system intelligently classifies tasks and routes simple queries to Gemini Flash ($0.0001/request) and complex reasoning tasks to Claude Sonnet 4 ($0.01/request). Requires GOOGLE_CLOUD_LOCATION=us-east5 for Claude model access. Future versions will expand this POC into a full enterprise model orchestration system with additional providers and more sophisticated routing logic.
Summary by CodeRabbit
New Features
Documentation