Skip to content

llm: fallback to legacy nearai.session key when loading DB session - #111

Merged
ilblackdragon merged 3 commits into
nearai:mainfrom
lawyered0:chore/nearai-session-fallback
Feb 17, 2026
Merged

ilblackdragon merged 3 commits into
nearai:mainfrom
lawyered0:chore/nearai-session-fallback

Conversation

@lawyered0

@lawyered0 lawyered0 commented Feb 16, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Addresses auth regressions for legacy installations by adding a backward-compatible read path when loading NEAR AI sessions from DB settings.

Change

  • In session loader, prefer nearai.session_token as primary key.
  • If missing, fall back to legacy nearai.session key.
  • Keep existing behavior unchanged when a valid session object is found.

Rationale

  • Some deployments may have sessions already persisted under the legacy key.
  • This avoids re-auth failures during upgrades without changing current write key.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @lawyered0, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a backward-compatible mechanism for loading NEAR AI sessions from the database. It aims to prevent authentication regressions for older installations by ensuring that sessions persisted under a legacy key can still be retrieved, thereby facilitating smoother upgrades without altering the current write key for new sessions.

Highlights

  • Session Loading Priority: The system now prioritizes loading the session using the nearai.session_token key.
  • Legacy Fallback: If nearai.session_token is not found, the system will attempt to load the session using the older nearai.session key for backward compatibility.
  • Existing Session Handling: The behavior remains unchanged if a valid session object is successfully retrieved, regardless of the key used.
Changelog
  • src/llm/session.rs
    • Implemented a conditional check to first attempt retrieving the session using nearai.session_token and then falling back to nearai.session if the primary key is not found.
    • Added a tracing::warn message when the fallback mechanism is engaged.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@lawyered0

Copy link
Copy Markdown
Contributor Author

Added as a follow-up split PR from PR #109 for clean audit scope.

Scope is limited to legacy-session compatibility in DB loader:

  • Primary read remains nearai.session_token
  • Fallback to nearai.session if primary key is missing

No other behavior changes included.

@lawyered0

Copy link
Copy Markdown
Contributor Author

This PR is intended as a fix for issue #108 (legacy session key migration path), scoped only to session read compatibility.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds a fallback mechanism to load legacy NEAR AI sessions from the database, which is a good improvement for backward compatibility. The implementation is correct, and I have a suggestion to improve the code's readability and make it more idiomatic.

Comment thread src/llm/session.rs Outdated
Comment on lines +431 to +455
let value = match store
.get_setting(&user_id, "nearai.session_token")
.await
.map_err(|e| LlmError::SessionRenewalFailed {
provider: "nearai".to_string(),
reason: format!("DB query failed: {}", e),
})?
.ok_or_else(|| LlmError::SessionRenewalFailed {
provider: "nearai".to_string(),
reason: "No session in DB".to_string(),
})?;
})? {
Some(value) => value,
None => {
tracing::warn!(
"nearai.session_token missing; falling back to legacy nearai.session for backwards compatibility"
);
store
.get_setting(&user_id, "nearai.session")
.await
.map_err(|e| LlmError::SessionRenewalFailed {
provider: "nearai".to_string(),
reason: format!("DB query failed: {}", e),
})?
.ok_or_else(|| LlmError::SessionRenewalFailed {
provider: "nearai".to_string(),
reason: "No session in DB".to_string(),
})?
}
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The match statement can be expressed more idiomatically as an if let expression for this case. Additionally, ok_or_else can be simplified to ok_or since the error value is not computed dynamically.

        let value = if let Some(value) = store
            .get_setting(&user_id, "nearai.session_token")
            .await
            .map_err(|e| LlmError::SessionRenewalFailed {
                provider: "nearai".to_string(),
                reason: format!("DB query failed: {}", e),
            })? {
            value
        } else {
            tracing::warn!(
                "nearai.session_token missing; falling back to legacy nearai.session for backwards compatibility"
            );
            store
                .get_setting(&user_id, "nearai.session")
                .await
                .map_err(|e| LlmError::SessionRenewalFailed {
                    provider: "nearai".to_string(),
                    reason: format!("DB query failed: {}", e),
                })?
                .ok_or(LlmError::SessionRenewalFailed {
                    provider: "nearai".to_string(),
                    reason: "No session in DB".to_string(),
                })?
        };

@lawyered0

Copy link
Copy Markdown
Contributor Author

Implemented Gemini’s style suggestion on the code path (if-let + ok_or form) in commit a21d271. No behavior change; just readability/idiomatic cleanup. This should clear that review nit.

@ilblackdragon ilblackdragon left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Clean, well-scoped change for backward compatibility.

What I checked:

  1. Correctness: The fallback logic is sound. When nearai.session_token returns None, it tries nearai.session before failing. The write path (save_session) still writes to nearai.session_token only (line 408), so after a successful save-then-load cycle the legacy key becomes unused -- which is the right migration behavior.

  2. Error handling: Both DB query errors and missing-key errors are handled identically between the primary and fallback paths. The map_err + ok_or chain is consistent with the rest of the file. Using ok_or instead of ok_or_else is fine here since the error construction is cheap (just two String allocations).

  3. Format assumption: Both keys feed into the same serde_json::from_value::<SessionData>() call (line 443-447). This assumes the legacy nearai.session key stores the same SessionData JSON shape (session_token, created_at, auth_provider). If any legacy deployment stored a different format (e.g., a bare token string), this would fail at deserialization rather than silently misbehaving -- which is the correct failure mode since the error message says "Failed to parse DB session".

  4. Code style: The if let form adopted in the second commit is idiomatic and reads well. The tracing::warn! log message clearly identifies what happened and why, which will help with debugging in production.

  5. No regressions: The write path is untouched, so current deployments continue writing to nearai.session_token. The fallback is read-only and only triggers when the primary key is absent.

Minor observation (non-blocking): The duplicated map_err closure for DB query failures appears in both the primary and fallback paths. A small helper like fn db_query_err(e: impl Display) -> LlmError could reduce repetition, but that is a broader pattern across this file (see lines 186, 231, 261, 320, etc.) and not worth addressing in this PR alone.

@tribendu

Copy link
Copy Markdown

Batch 2 PR Review: nearai/ironclaw PRs 111, 110, 109, 103, 95, 74

Reviewer: AI Sub-Agent
Date: February 17, 2026
Review Method: Manual diff analysis with focus on code quality, bugs, security, performance, test coverage, and Rust best practices


PR #111: Fix backwards compatibility for nearai.session_token

Summary

This PR adds backwards compatibility for the nearai session management by implementing a fallback mechanism. When nearai.session_token is not found, the system now attempts to fall back to the legacy nearai.session setting. The change is minimal and focused on the session token retrieval logic in SessionManager::renew_session(). The implementation uses a graceful fallback pattern with proper error handling and logging.

Pros

  • Graceful degradation: Implements backwards compatibility without breaking existing functionality
  • Clear error handling: Preserves original error handling for both primary and fallback paths
  • Informative logging: Adds a warning message when falling back to legacy session storage
  • Minimal change scope: Only affects session token retrieval logic, reducing risk of regression
  • Proper pattern: Uses idiomatic Rust if let Some(...) for optional value handling

Concerns

  • Technical debt: This adds maintenance burden by keeping two parallel session storage schemas
  • No migration path: There's no clear deprecation timeline or migration strategy for the legacy schema
  • Silent fallback: The warning log may be missed in production, leading to continued use of deprecated schema
  • Error propagation: The error messages remain generic - they could be more specific about whether the primary or fallback path failed

Suggestions

  • Add a migration function to convert legacy sessions to the new format periodically
  • Consider adding telemetry/metrics to track usage of the legacy fallback path
  • Document the expected lifecycle and deprecation timeline for nearai.session
  • Add a configuration option to disable the fallback after migration is complete
  • Consider making the fallback opt-in rather than silent to encourage migration

PR #110: Add env docs for local LLM providers

Summary

This PR updates the .env.example file to document environment variables for local LLM providers including Ollama, LM Studio, vLLM, and other OpenAI-compatible endpoints. The changes are purely documentation-focused, providing clear examples of environment variable configurations for users who want to use local models instead of the default NEAR AI backend.

Pros

  • User-friendly: Makes it easy for new users to configure local LLM providers
  • Comprehensive: Covers multiple backend options (Ollama, OpenAI-compatible)
  • Clear defaults: Shows default values for URLs and ports
  • Well-organized: Groups related configurations logically
  • No code changes: Documentation-only PR, zero risk of breaking functionality

Concerns

  • No validation: Environment variables are documented but there's no validation that the backend value matches the selected provider
  • Incomplete examples: Example model names are specific but may not match what users have installed
  • Missing authentication notes: Doesn't document that LLM_API_KEY is optional for local servers

Suggestions

  • Add comments explaining when LLM_API_KEY is optional vs required
  • Include a link to documentation for each provider
  • Add a note about model availability requiring local installation
  • Consider adding environment variable validation logic in a future PR
  • Document how to switch between providers at runtime
  • Add example commands to test connections (e.g., curl http://localhost:11434/api/tags)

PR #109: Normalize memory search query and update marked.js

Summary

This PR addresses two security and stability issues in the web interface. First, it adds input normalization for memory search queries to prevent excessive query lengths and invalid input types. Second, it updates the marked.js dependency to a specific version with integrity hashing for supply chain security. The changes are defensive in nature, preventing potential DoS attacks and ensuring the integrity of third-party JavaScript dependencies.

Pros

  • Input validation: Adds MEMORY_SEARCH_QUERY_MAX_LENGTH constant (100 chars) to prevent excessively long queries
  • Type safety: Normalizes queries to string type, preventing crashes from non-string input
  • Supply chain security: Uses SRI (Subresource Integrity) hashing for marked.js CDN dependency
  • Consistent application: Applies normalization throughout search functions (searchMemory, snippetAround, highlightQuery)
  • Clear error handling: Early return when normalized query is empty

Concerns

  • Arbitrary limit: 100-character limit may be too restrictive for complex semantic search queries
  • Silent truncation: Queries longer than the limit are silently truncated without user feedback
  • CDN dependency: Still relies on external CDN for marked.js despite SRI - could use bundled version
  • No rate limiting: Query length limit doesn't prevent rapid-fire spam queries
  • Client-side only: Validation happens on client side only - server should also validate

Suggestions

  • Increase the limit or make it configurable (e.g., 200-500 characters)
  • Add user feedback when query is truncated (toast message or inline warning)
  • Implement rate limiting on the /api/memory/search endpoint
  • Add server-side query validation to defend against bypassed client checks
  • Consider bundling marked.js with the application to eliminate CDN dependency
  • Add unit tests for the normalization function covering edge cases (null, undefined, empty string, very long string)

PR #103: Per-request model override for OpenAI-compatible API

Summary

This is a significant feature PR that adds per-request model override capability across the entire LLM provider ecosystem. Previously, all requests used the active model, but now clients can specify a different model per request. The changes span multiple modules: request structs now include optional model fields, providers check for request-level models before falling back to defaults, OpenAI-compatible endpoint validates model names, and comprehensive integration tests verify functionality. The PR includes migration documentation and removes the previous strict model validation that returned 404 for non-active models.

Pros

  • Well-architected: Uses Option with builder pattern (with_model()) for clean API
  • Comprehensive testing: Updated integration tests to mock model tracking and verify model override works
  • Full stack coverage: Changes propagate from API entry points through worker layer to actual providers
  • Good error handling: Validates model name length (MAX_MODEL_NAME_BYTES: 256) before processing
  • Backwards compatible: Existing behavior preserved when model field is not provided
  • Code quality: Consistent pattern across CompletionRequest and ToolCompletionRequest
  • Documentation: Updates FEATURE_PARITY.md to reflect new capability

Concerns

  • No model validation: The endpoint now accepts ANY model name without checking if it exists or is configured
  • Validation bypasses providers: Validation happens at HTTP layer, but providers don't validate model availability
  • Security implications: Users can request models they shouldn't have access to; cost tracking may break
  • Incomplete fallback: If provider doesn't support the requested model, behavior is undefined
  • Missing permissions: No RBAC or policy checking for model override capability
  • Potential for confusion: Active model concept remains but can be overridden per request

Suggestions

  • Add a method to validate if a model exists/can be used before making the request
  • Implement optional model allowlist/denylist configuration per user or API key
  • Add telemetry to track model usage patterns across different requests
  • Document the precedence: request model > active model > provider default
  • Consider adding a force_model flag that fails fast if model is unavailable
  • Add tests for error cases (non-existent model, invalid characters, empty string)
  • Update cost tracking to handle per-request model pricing differences
  • Consider adding model availability check to LlmProvider trait's set_model() method

PR #95: Add Venice AI provider and embeddings

Summary

This PR adds comprehensive support for Venice AI as both an LLM provider and an embeddings provider. It introduces a new VeniceProvider module with full API integration including Venice-specific features like web search, web scraping, and dynamic model pricing fetched from the /models endpoint. The PR also adds VeniceEmbeddings for semantic search, updates configuration handling, and integrates Venice into the setup wizard. The implementation includes extensive unit tests, proper error handling, caching of model catalogs with TTL, and follows the existing provider patterns.

Pros

  • Fully featured: Implements complete Venice API including proprietary parameters (web search, scraping, system prompt)
  • Dynamic pricing: Fetches model catalog with pricing from Venice API, updates cost tracking automatically
  • Good caching: Uses 1-hour TTL for model catalog to reduce API calls
  • Comprehensive testing: Extensive unit tests for message conversion, serialization, cost calculation
  • Type safety:Uses properly typed structures and enums for Venice-specific functionality
  • Consistent patterns: Follows existing provider patterns (NearAiProvider, etc.)
  • Error handling: Proper HTTP error mapping (401→AuthFailed, 429→RateLimited)
  • Config validation: Validates Venice-specific config values (web_search must be "off"/"on"/"auto")

Concerns

  • Large file: venice.rs is 880 lines - consider splitting into smaller modules
  • Lock choice: Uses std::sync::RwLock with comment about async safety - risk if async code is added later
  • Secret handling: API key exposed as String in multiple places despite being wrapped in SecretString in Config
  • No retry logic: Network requests don't have exponential backoff for transient failures
  • Hardcoded timeout: 120-second timeout may be too long or too short for different use cases
  • Cache misses on errors: If catalog fetch fails, stale cache is kept indefinitely
  • Embeddings dimension hardcoding: Magic numbers for dimensions (1536, 3072) scattered in code

Suggestions

  • Split venice.rs into multiple modules: provider.rs, models.rs, api.rs
  • Consider using tokio::sync::RwLock for future-proofing with async code
  • Add retry logic with exponential backoff for API requests
  • Make timeout configurable via VeniceConfig
  • Add metrics for cache hit/miss ratio
  • Document the security model: when are secrets exposed?
  • Extract dimension mapping to a constant or helper function
  • Add integration test that calls actual Venice API with test credentials
  • Consider adding health check endpoint for provider connectivity
  • Add support for streaming responses (if Venice API supports it)

PR #74: Security fix: Enhanced HTTP response size validation

Summary

This PR strengthens the HTTP tool's defense against OOM attacks by implementing two-stage response size validation. First, it checks the Content-Length header before downloading any content to immediately reject oversized responses. Second, it streams the response body with a hard size cap during download, protecting against malicious or misconfigured servers that may send incorrect or missing Content-Length headers. The implementation uses streaming with futures::StreamExt and enforces a consistent 5 MB limit across both stages.

Pros

  • Defense in depth: Two-stage validation (header check + streaming cap) protects against multiple attack vectors
  • Early rejection: Aborts request before downloading potentially malicious content
  • Streaming approach: Memory-efficient, doesn't load entire response before checking size
  • Informative logging: Warns with URL and size details when rejecting responses
  • Consistent constant: Uses same 5 MB limit as WASM HTTP wrapper
  • Well-documented: Clear comments explain the security rationale and 5 MB justification
  • Test coverage: Includes test verifying the constant value is reasonable

Concerns

  • Header spoofing: Relies on server sending accurate Content-Length - malicious servers can send small value then send large body
  • No byte limit on individual chunks: Only checks cumulative size; individual chunks could still be large
  • Error message exposure: Returns detailed size information in errors which might leak information
  • No rate limiting: An attacker could still send many small requests to exhaust resources
  • Hardcoded limit: 5 MB may be too small for some legitimate use cases (e.g., downloading large JSON datasets)

Suggestions

  • Add per-chunk size limit (e.g., 1 MB max per chunk) in addition to cumulative limit
  • Consider making MAX_RESPONSE_SIZE configurable per tool or per request
  • Add telemetry/metrics for rejected responses to detect DDoS patterns
  • Implement request rate limiting in addition to size limiting
  • Add option to override limit for authorized users/admins
  • Consider adding a timeout for the streaming operation
  • Add test cases for Content-Length header spoofing attempt
  • Document how users can work around the limit if needed (e.g., using multiple smaller requests with pagination)

Overall Recommendations

High Priority

  1. PR feat: support per-request model override in /v1/chat/completions #103: Address the security implications of unrestricted model override before merging - add allowlist/denylist functionality
  2. PR fix: check Content-Length before downloading HTTP response body #74: Consider making the size limit configurable for flexibility while maintaining security

Medium Priority

  1. PR feat: add Venice AI as first-class LLM backend #95: Refactor the large venice.rs file and add retry logic for network resilience
  2. PR web: add integrity check for marked CDN and cap highlight regex input #109: Add server-side validation to complement client-side checks
  3. PR llm: fallback to legacy nearai.session key when loading DB session #111: Add migration strategy for legacy session tokens

Low Priority

  1. PR docs: add .env.example examples for Ollama and OpenAI-compatible #110: Add validation documentation and testing commands
  2. PR feat: add Venice AI as first-class LLM backend #95: Extract magic numbers and add integration tests
  3. PR web: add integrity check for marked CDN and cap highlight regex input #109: Consider removing CDN dependency for marked.js

General Observations

  • All PRs follow Rust best practices and maintain code quality
  • Test coverage is generally good, though some integration tests could be expanded
  • Error handling is consistent across the codebase
  • Documentation (comments and inline docs) is clear and helpful
  • Most PRs are backwards compatible, which is good for production deployments

@ilblackdragon
ilblackdragon merged commit 956037c into nearai:main Feb 17, 2026
@github-actions github-actions Bot mentioned this pull request Feb 17, 2026
jaswinder6991 pushed a commit to jaswinder6991/ironclaw that referenced this pull request Feb 26, 2026
…earai#111)

* llm: fallback to legacy nearai.session when loading DB session

* llm: simplify session fallback load with if-let form

---------

Co-authored-by: Clawyered <clawyered@macbookair.home>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
…earai#111)

* llm: fallback to legacy nearai.session when loading DB session

* llm: simplify session fallback load with if-let form

---------

Co-authored-by: Clawyered <clawyered@macbookair.home>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants