Repository navigation
feat(memory): Implement token based summarization - #701
Yaswanth-2874 wants to merge 1 commit into
Conversation
|
Important Review skippedAuto incremental reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the You can disable this status message by setting the WalkthroughTransition conversation memory from turn-count to token-threshold management: introduce token-based summarization, add provider/model metadata, change storeConversationTurn to an options object, extend message/session types and Redis config, and wire summarization through NeuroLink and Redis-backed managers. Changes
Sequence Diagram(s)sequenceDiagram
participant Client
participant NeuroLink
participant MemoryMgr as ConversationMemoryManager
participant TokenUtils
participant Redis
participant Summarizer as Summarization Service
Client->>NeuroLink: generate/stream request (model)
NeuroLink->>NeuroLink: extract providerDetails, enableSummarization
NeuroLink->>MemoryMgr: storeConversationTurn(options)
MemoryMgr->>MemoryMgr: validateAndPrepareMessage (may truncate)
MemoryMgr->>TokenUtils: estimateTokens(messages)
TokenUtils-->>MemoryMgr: token count
MemoryMgr->>MemoryMgr: getEffectiveTokenThreshold
MemoryMgr->>MemoryMgr: checkAndSummarize(session)
alt summarization required
MemoryMgr->>MemoryMgr: findSplitIndexByTokens -> select recent messages
MemoryMgr->>Summarizer: generateSummary(recent messages, prompt)
Summarizer-->>MemoryMgr: summary text
MemoryMgr->>MemoryMgr: createSummarySystemMessage(metadata)
MemoryMgr->>Redis: persist summarizedUpToMessageId + summarizedMessage
else no summarization
MemoryMgr->>Redis: persist new turn
end
Redis-->>MemoryMgr: ack
MemoryMgr-->>NeuroLink: ack
NeuroLink-->>Client: final response
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes
Possibly related PRs
Suggested reviewers
Poem
Pre-merge checks and finishing touches✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/lib/neurolink.ts (1)
2772-2813: Add explicit string type check foruserIdin streaming memory write pathThe streaming completion handler correctly accumulates
aiResponseand writes a unified turn viaconversationMemory.storeConversationTurn, includingproviderDetailsandenableSummarization. One type-safety improvement:
userIdis cast directly viaas stringwithout validation. The non-streaming path uses an explicit guard (typeof context.userId === "string" ? context.userId : undefined), which is safer and more consistent. Align the stream path to match:- const userId = ( - enhancedOptions.context as Record<string, unknown> - )?.userId as string; + const rawUserId = (enhancedOptions.context as Record<string, unknown>) + ?.userId; + const userId = + typeof rawUserId === "string" ? rawUserId : undefined;
🧹 Nitpick comments (14)
src/lib/utils/conversationMemoryUtils.ts (1)
129-170: Redis availability check is well-implemented.The function properly:
- Uses short timeout (5000ms) appropriate for availability probing
- Limits retries to 1 to fail fast
- Ensures cleanup in the
finallyblock with error handling for quit failuresOne minor consideration: the
testClientvariable initialization tonulland subsequent type narrowing works, but you could use optional chaining in the finally block for slightly cleaner code.Optional: Simplify cleanup with optional chaining
} finally { - if (testClient) { - try { - await testClient.quit(); - logger.debug("Redis test client disconnected successfully"); - } catch (quitError) { - logger.debug("Error during Redis test client disconnect", { - error: - quitError instanceof Error ? quitError.message : String(quitError), - }); - } + try { + await testClient?.quit(); + if (testClient) logger.debug("Redis test client disconnected successfully"); + } catch (quitError) { + logger.debug("Error during Redis test client disconnect", { + error: + quitError instanceof Error ? quitError.message : String(quitError), + }); } }src/lib/utils/conversationMemory.ts (3)
230-272: Align summary message metadata withsummarizesFromfor better traceability
buildContextFromPointercorrectly injects a synthetic summaryChatMessageand setsmetadata.isSummaryandmetadata.summarizesTo, but leavessummarizesFromunset. For downstream tooling that wants to know the exact covered range, filling both endpoints makes this more useful and consistent withcreateSummarySystemMessagein the managers.Optional metadata enhancement
const summaryMessage: ChatMessage = { id: `summary-${session.summarizedUpToMessageId}`, role: "system", content: `Previous conversation summary: ${session.summarizedMessage}`, timestamp: new Date().toISOString(), metadata: { isSummary: true, - summarizesTo: session.summarizedUpToMessageId, + summarizesFrom: session.messages[0]?.id, + summarizesTo: session.summarizedUpToMessageId, }, };
317-375: Clarify / simplify error handling in token-threshold helpers
calculateTokenThresholdalready catches and logs errors and falls back toDEFAULT_FALLBACK_THRESHOLD, so thetry/catchingetEffectiveTokenThresholdwill never see an exception from it. The outertry/catchis therefore redundant and slightly obscures the priority logic.You could simplify
getEffectiveTokenThresholdto rely oncalculateTokenThreshold’s own fallback and keep the priority comments as-is.Suggested simplification
export function getEffectiveTokenThreshold( provider: string, model: string, envOverride?: number, sessionOverride?: number, ): number { // Priority 1: Session-level override if (sessionOverride && sessionOverride > 0) { return sessionOverride; } // Priority 2: Environment variable override if (envOverride && envOverride > 0) { return envOverride; } - // Priority 3: Model-based calculation (80% of context window) - try { - return calculateTokenThreshold(provider, model); - } catch (error) { - logger.warn("Failed to calculate effective threshold, using fallback", { - provider, - model, - error: error instanceof Error ? error.message : String(error), - }); - // Priority 4: Fallback for unknown models - return DEFAULT_FALLBACK_THRESHOLD; - } + // Priority 3: Model-based calculation (80% of context window) with internal fallback + return calculateTokenThreshold(provider, model); }
386-417: Avoid tight coupling and potential cycles by lazily importingNeuroLinkingenerateSummaryThis utility module is now pulling in the full
NeuroLinkclass via a static import and using it only insidegenerateSummary. GivenNeuroLinkalready imports this file forstoreConversationTurn/getConversationMessages, this creates a bidirectional dependency and also forces all of NeuroLink’s heavy initialization code to be eagerly linked whenever these helpers are imported.You can decouple the layers and make initialization cheaper by lazily importing
NeuroLinkinsidegenerateSummary:Proposed lazy import for `NeuroLink`
-import { NeuroLink } from "../neurolink.js"; @@ export async function generateSummary( messages: ChatMessage[], config: ConversationMemoryConfig, logPrefix = "[ConversationMemory]", previousSummary?: string, ): Promise<string | null> { const summarizationPrompt = createSummarizationPrompt( messages, previousSummary, ); - const summarizer = new NeuroLink({ - conversationMemory: { enabled: false }, - }); + // Lazy‑load NeuroLink to avoid hard module cycles and unnecessary startup work + const { NeuroLink } = await import("../neurolink.js"); + const summarizer = new NeuroLink({ + conversationMemory: { enabled: false }, + });This keeps the API the same but breaks the static coupling to the main SDK class.
src/lib/neurolink.ts (1)
3090-3129: Fallback stream memory write reuses context defensively but could be simplifiedIn the fallback streaming path, session and user IDs are computed as:
const sessionId = (enhancedOptions?.context as Record<string, unknown>)?.sessionId as string; const userId = (enhancedOptions?.context as Record<string, unknown>)?.userId as string; await self.conversationMemory.storeConversationTurn({ sessionId: sessionId || (options.context?.sessionId as string), userId: userId || (options.context?.userId as string), ... });Given the outer
ifalready checksenhancedOptions?.context?.sessionId,sessionIdshould always be a truthy string, so the|| options.context?.sessionIdfallback is effectively dead code. The same applies touserIdif you add a proper type guard as suggested for the main stream path.You can simplify this block and avoid redundant
as stringcasts, which makes it easier to reason about and avoids accidentally passingundefinedintoStoreConversationTurnOptionsin edge cases.src/lib/types/conversation.ts (2)
39-47: Deprecated turn-based config fields are still read in logs
maxTurnsPerSession,summarizationThresholdTurns, andsummarizationTargetTurnsare now marked@deprecatedin favor oftokenThreshold, but the Redis and in-memory managers still logmaxTurnsPerSessionin their initialization messages.That’s harmless but slightly confusing for users migrating to token-based configs. Consider either:
- Updating log messages to mention
tokenThreshold(and possibly the effective threshold), or- Clearly flagging these fields as ignored in the runtime logic when
tokenThresholdis set.
250-261: Consider carryingtokenThresholdintoStoreConversationTurnOptionswhen using per-session overrides
StoreConversationTurnOptionscurrently carries sessionId, userId, messages, timestamps, providerDetails, andenableSummarization. The managers then recompute an effective threshold viagetEffectiveTokenThreshold, usingsession.tokenThresholdas a possible override.Right now there’s no way for a caller to set
session.tokenThresholdthrough this options type; the managers never write it, they only read it. If per-session overrides are intended, you may want to:
- Add an optional
tokenThreshold?: numbertoStoreConversationTurnOptions, and- Persist it into the session/Redis object on first write.
Otherwise the
sessionOverridepath ingetEffectiveTokenThresholdwill effectively never be used.src/lib/core/redisConversationMemoryManager.ts (3)
328-425: Token-threshold computation andenableSummarizationprecedence look correct, butconversation.tokenThresholdis never persistedIn
storeConversationTurnyou:
- Compute
tokenThresholdviagetEffectiveTokenThresholdwhenoptions.providerDetailsis present, otherwise fall back tothis.config.tokenThreshold || 50000.- Use per-request
options.enableSummarization(if defined) to override the instance-levelconfig.enableSummarization.That precedence is sensible. However,
conversation.tokenThresholdis passed intogetEffectiveTokenThresholdas a session override and later re-exposed in theSessionMemoryshim, but it is never actually set on theconversationobject in this file. As a result, the “sessionOverride” branch ingetEffectiveTokenThresholdis effectively dead for Redis-backed sessions.If you intend to support per-session thresholds, consider persisting
tokenThresholdon first calculation:Suggested persistence of `conversation.tokenThreshold`
- const tokenThreshold = options.providerDetails + let tokenThreshold = options.providerDetails ? getEffectiveTokenThreshold( options.providerDetails.provider, options.providerDetails.model, this.config.tokenThreshold, conversation.tokenThreshold, ) : this.config.tokenThreshold || 50000; + + // Persist effective threshold on the conversation so future turns can reuse it + conversation.tokenThreshold = conversation.tokenThreshold ?? tokenThreshold;
453-475: Background summarization runs with a captured snapshot; be aware of concurrent write semanticsThe
setImmediatecallback callscheckAndSummarize(conversation, tokenThreshold, options.sessionId, options.userId), whereconversationis the deserialized object from this particularstoreConversationTurncall. If multiple turns for the same session are stored in quick succession, each call will:
- Read its own snapshot of the conversation from Redis.
- Enqueue a summarization job based on that snapshot.
- Write its snapshot back to Redis (in
summarizeSessionTokenBased).Because
setImmediateis FIFO, the last turn’s summarization job should generally win, but intermediate summarization jobs can overwrite newer fields (e.g., pointer or token counts) computed by later jobs if anything unusual happens in scheduling.Given this is already a best‑effort background process and not part of the main request path, this is probably acceptable, but it’s worth noting that the current implementation does not guarantee strictly monotonic updates to
summarizedUpToMessageId/summarizedMessageunder high concurrency for the same session.If you see inconsistent summaries in practice, a minimal mitigation would be to re‑fetch the latest conversation inside
checkAndSummarizeusingsessionId/userIdbefore computing and writing summary data, instead of relying solely on the capturedconversationobject.
653-709: Context building and tool-message filtering semantics differ from in-memory manager
buildContextMessages:
- Deserializes the Redis conversation,
- Wraps it as
SessionMemory,- Uses
buildContextFromPointerto inject the summary and slice from the pointer,- Optionally filters out
tool_call/tool_resultmessages whenconfig.enableSummarization === true.In the in-memory
ConversationMemoryManager.buildContextMessages, context is also built viabuildContextFromPointer, but no role-based filtering is applied, so tool messages are always included.If the goal is to have symmetrical behavior between Redis and in-memory memory (especially when summarization is enabled), consider aligning the in-memory implementation to apply the same tool-message filtering. Right now, behavior differs by storage backend.
src/lib/core/conversationMemoryManager.ts (4)
65-103: Per-session threshold override is computed but not persisted
storeConversationTurncomputestokenThresholdviagetEffectiveTokenThreshold, supplyingsession.tokenThresholdas a possible override, butcreateNewSessionnever setstokenThresholdand this method doesn’t update it either. That means thesessionOverridebranch ingetEffectiveTokenThresholdis effectively unused for in-memory sessions.If you want per-session thresholds to stick after the first effective calculation, mirror the Redis suggestion and persist it onto the session:
Example persistence on `SessionMemory`
- const tokenThreshold = options.providerDetails + let tokenThreshold = options.providerDetails ? getEffectiveTokenThreshold( options.providerDetails.provider, options.providerDetails.model, this.config.tokenThreshold, session.tokenThreshold, ) : this.config.tokenThreshold || 50000; + + session.tokenThreshold = session.tokenThreshold ?? tokenThreshold;
130-174:validateAndPrepareMessagetruncation logic is correct; async modifier is unnecessaryThe helper:
- Estimates tokens with
TokenUtils.estimateTokenCount,- Truncates to
threshold * MEMORY_THRESHOLD_PERCENTAGEwhen necessary,- Marks truncated messages via
metadata.truncated = true,- Always assigns an
idandtimestamp.This matches the new ChatMessage shape and keeps overlong messages in check. The function is marked
asyncbut doesn’tawaitanything, so it always returns a resolvedPromise<ChatMessage>; that’s harmless but adds a bit of noise. You could safely make it synchronous and update the call sites to dropawaitfor slightly simpler code.
217-225: In-memorybuildContextMessagesshould probably mirror Redis tool-message filteringThe in-memory
buildContextMessagesnow simply returnsbuildContextFromPointer(session)without any additional filtering, while the Redis manager optionally removestool_call/tool_resultmessages when summarization is enabled. That means the same logical session can produce different contexts depending on backend.If you want parity, consider applying the same optional role-based filter here (perhaps driven by
this.config.enableSummarization) so callers get consistent behavior regardless of storage implementation.
252-304: Token-based summarization flow is correct but currently includes tool messages
summarizeSessionTokenBased:
- Starts from
summarizedUpToMessageId+ 1,- Slices
recentMessages,- Uses token-based splitting via
findSplitIndexByTokenswithRECENT_MESSAGES_RATIO,- Calls
generateSummarywith the new chunk and previous summary,- Advances the pointer and stores the concatenated summary.
Unlike the Redis implementation, it does not filter out
tool_call/tool_resultmessages before summarizing, so tool traffic can dominate the token budget in memory-backed mode.If that’s not intentional, you could mirror the Redis behavior by filtering roles before the split:
Optional alignment with Redis summarization
- const recentMessages = session.messages.slice(startIndex); + const recentMessages = session.messages.slice(startIndex); + const filteredRecentMessages = recentMessages.filter( + (msg) => msg.role !== "tool_call" && msg.role !== "tool_result", + ); @@ - const targetRecentTokens = threshold * RECENT_MESSAGES_RATIO; - const splitIndex = await this.findSplitIndexByTokens( - recentMessages, - targetRecentTokens, - ); - const messagesToSummarize = recentMessages.slice(0, splitIndex); + const targetRecentTokens = threshold * RECENT_MESSAGES_RATIO; + const splitIndex = await this.findSplitIndexByTokens( + filteredRecentMessages, + targetRecentTokens, + ); + const messagesToSummarize = filteredRecentMessages.slice(0, splitIndex);
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (14)
.env.example(2 hunks)src/lib/config/conversationMemory.ts(1 hunks)src/lib/core/conversationMemoryFactory.ts(0 hunks)src/lib/core/conversationMemoryInitializer.ts(0 hunks)src/lib/core/conversationMemoryManager.ts(2 hunks)src/lib/core/redisConversationMemoryManager.ts(10 hunks)src/lib/neurolink.ts(4 hunks)src/lib/types/conversation.ts(8 hunks)src/lib/types/generateTypes.ts(1 hunks)src/lib/types/sdkTypes.ts(0 hunks)src/lib/types/streamTypes.ts(1 hunks)src/lib/utils/conversationMemory.ts(3 hunks)src/lib/utils/conversationMemoryUtils.ts(2 hunks)src/lib/utils/redis.ts(0 hunks)
💤 Files with no reviewable changes (4)
- src/lib/types/sdkTypes.ts
- src/lib/utils/redis.ts
- src/lib/core/conversationMemoryInitializer.ts
- src/lib/core/conversationMemoryFactory.ts
🧰 Additional context used
📓 Path-based instructions (2)
**/*.{ts,tsx}
📄 CodeRabbit inference engine (CLAUDE.md)
**/*.{ts,tsx}: Maintain strict TypeScript type safety across all modules with comprehensive type definitions organized by domain to avoid circular dependencies
Use ErrorFactory for creating typed errors throughout the application
Wrap async operations with withTimeout utility for timeout handling
Files:
src/lib/types/streamTypes.tssrc/lib/config/conversationMemory.tssrc/lib/utils/conversationMemory.tssrc/lib/utils/conversationMemoryUtils.tssrc/lib/core/conversationMemoryManager.tssrc/lib/core/redisConversationMemoryManager.tssrc/lib/types/generateTypes.tssrc/lib/types/conversation.tssrc/lib/neurolink.ts
**/types/**/*.ts
📄 CodeRabbit inference engine (CLAUDE.md)
Type definitions must be organized by domain (providers, generation, streaming, MCP, etc.) to avoid circular dependencies
Files:
src/lib/types/streamTypes.tssrc/lib/types/generateTypes.tssrc/lib/types/conversation.ts
🧠 Learnings (14)
📓 Common learnings
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Memory management should use Redis for distributed memory in production and in-memory store for development, with conversation summarization for long contexts
Learnt from: BoraYaswanthReddy
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-12-14T18:33:26.766Z
Learning: In the neurolink repository's Redis conversation memory implementation, message IDs were changed from sequential integers to UUIDs (using `generateUniqueId()`) as a security improvement to prevent enumeration attacks and information leakage about conversation volumes and patterns.
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Memory management should use Redis for distributed memory in production and in-memory store for development, with conversation summarization for long contexts
Applied to files:
.env.examplesrc/lib/config/conversationMemory.tssrc/lib/utils/conversationMemory.tssrc/lib/core/conversationMemoryManager.tssrc/lib/core/redisConversationMemoryManager.tssrc/lib/types/conversation.ts
📚 Learning: 2025-12-14T18:33:26.766Z
Learnt from: BoraYaswanthReddy
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-12-14T18:33:26.766Z
Learning: In the neurolink repository's Redis conversation memory implementation, message IDs were changed from sequential integers to UUIDs (using `generateUniqueId()`) as a security improvement to prevent enumeration attacks and information leakage about conversation volumes and patterns.
Applied to files:
.env.examplesrc/lib/core/redisConversationMemoryManager.tssrc/lib/neurolink.ts
📚 Learning: 2025-12-14T18:33:26.766Z
Learnt from: BoraYaswanthReddy
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-12-14T18:33:26.766Z
Learning: In the neurolink repository, the `separateLLMContext` configuration exists primarily to support agentic loops that need full conversation context including tool messages. The default is intentionally `true` because separating tool messages from LLM context is considered the better default behavior. For CLI usage, separation is always enabled and the option is not exposed to users.
Applied to files:
.env.example
📚 Learning: 2025-09-01T22:58:39.149Z
Learnt from: sudharsan-juspay
Repo: juspay/neurolink PR: 140
File: src/lib/core/types.ts:198-203
Timestamp: 2025-09-01T22:58:39.149Z
Learning: In src/lib/core/types.ts, StreamOptions (imported from streamTypes.js) and StreamingOptions are intentionally different types with different use cases. StreamingOptions is for unified AI requests with multiple provider configurations, while StreamOptions is for individual streaming operations.
Applied to files:
src/lib/types/streamTypes.ts
📚 Learning: 2025-09-01T06:15:59.759Z
Learnt from: amreetkhuntia
Repo: juspay/neurolink PR: 133
File: src/lib/core/types.ts:208-210
Timestamp: 2025-09-01T06:15:59.759Z
Learning: The middleware?: MiddlewareFactoryOptions field is already present in both TextGenerationOptions and StreamOptions interfaces in the neurolink codebase.
Applied to files:
src/lib/types/streamTypes.ts
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/cli/loop/session.ts : Loop mode interactive sessions should be implemented in src/cli/loop/session.ts with persistent conversation memory and session-wide configuration
Applied to files:
src/lib/utils/conversationMemory.tssrc/lib/core/conversationMemoryManager.tssrc/lib/core/redisConversationMemoryManager.tssrc/lib/types/conversation.ts
📚 Learning: 2025-09-17T17:55:15.261Z
Learnt from: RajuSudhar
Repo: juspay/neurolink PR: 173
File: src/lib/index.ts:16-16
Timestamp: 2025-09-17T17:55:15.261Z
Learning: In src/lib/types/providers.ts, ProviderConfig was renamed to AIModelProviderConfig to deduplicate type names, as there was an existing ProviderConfig type that better suited the "ProviderConfig" name. This was an intentional breaking change for better type organization.
Applied to files:
src/lib/utils/conversationMemoryUtils.tssrc/lib/types/conversation.tssrc/lib/neurolink.ts
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/types/index.ts : Add new provider names to the AIProviderName enum in src/lib/types/index.ts when adding a new provider
Applied to files:
src/lib/utils/conversationMemoryUtils.tssrc/lib/types/conversation.ts
📚 Learning: 2025-09-17T18:14:34.960Z
Learnt from: RajuSudhar
Repo: juspay/neurolink PR: 173
File: src/lib/types/index.ts:58-62
Timestamp: 2025-09-17T18:14:34.960Z
Learning: RajuSudhar explained that in the Neurolink codebase, there are multiple ProviderConfig types causing inconsistency. One existing ProviderConfig type better suited the "ProviderConfig" name, so they renamed the less-suitable one to AIModelProviderConfig to free up the name. Adding backward compatibility aliases would worsen naming inconsistency rather than help. The remaining duplicates will be systematically deduplicated in the 07-Types-Module.md TODO as part of their phased refactor approach.
Applied to files:
src/lib/utils/conversationMemoryUtils.ts
📚 Learning: 2025-12-12T20:11:17.070Z
Learnt from: Yaswanth-2874
Repo: juspay/neurolink PR: 672
File: src/lib/core/redisConversationMemoryManager.ts:1082-1091
Timestamp: 2025-12-12T20:11:17.070Z
Learning: In the Redis conversation memory implementation, LLM context keys (when `separateLLMContext` is enabled) intentionally use only sessionId without userId: `llm:context:${sessionId}`. This is by design to scope LLM context purely at the session level, relying on sessionId global uniqueness.
Applied to files:
src/lib/core/redisConversationMemoryManager.ts
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/utils/messageBuilder.ts : Message construction must be handled through MessageBuilder in src/lib/utils/messageBuilder.ts, which handles text, images, PDFs, and CSV files with provider-specific adapters
Applied to files:
src/lib/types/conversation.ts
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/providers/*.ts : Providers must extend a base provider or implement the provider interface and register in ProviderRegistry.registerAllProviders() with provider name, factory function, default model, and aliases
Applied to files:
src/lib/types/conversation.ts
📚 Learning: 2025-09-24T06:42:06.088Z
Learnt from: amreetkhuntia
Repo: juspay/neurolink PR: 185
File: src/lib/evaluation/contextBuilder.ts:79-85
Timestamp: 2025-09-24T06:42:06.088Z
Learning: In the NeuroLink codebase, using `(options.prompt || [])` pattern for handling potentially undefined prompt arrays is the preferred approach over extracting to a normalized variable when building conversation history in the ContextBuilder class.
Applied to files:
src/lib/neurolink.ts
🧬 Code graph analysis (5)
src/lib/config/conversationMemory.ts (1)
src/lib/types/conversation.ts (1)
ConversationMemoryConfig(11-47)
src/lib/utils/conversationMemory.ts (4)
src/lib/types/conversation.ts (4)
ProviderDetails(397-400)SessionMemory(52-97)ChatMessage(113-152)ConversationMemoryConfig(11-47)src/lib/types/sdkTypes.ts (3)
SessionMemory(204-204)ChatMessage(205-205)ConversationMemoryConfig(203-203)src/lib/config/conversationMemory.ts (2)
MEMORY_THRESHOLD_PERCENTAGE(40-40)DEFAULT_FALLBACK_THRESHOLD(45-45)src/lib/neurolink.ts (1)
NeuroLink(152-5954)
src/lib/utils/conversationMemoryUtils.ts (1)
src/lib/types/conversation.ts (1)
ProviderDetails(397-400)
src/lib/core/conversationMemoryManager.ts (4)
src/lib/types/conversation.ts (4)
StoreConversationTurnOptions(253-261)ConversationMemoryError(212-225)ChatMessage(113-152)SessionMemory(52-97)src/lib/utils/conversationMemory.ts (3)
getEffectiveTokenThreshold(347-375)buildContextFromPointer(230-272)generateSummary(386-418)src/lib/types/sdkTypes.ts (3)
ConversationMemoryError(209-209)ChatMessage(205-205)SessionMemory(204-204)src/lib/config/conversationMemory.ts (2)
MEMORY_THRESHOLD_PERCENTAGE(40-40)RECENT_MESSAGES_RATIO(52-52)
src/lib/neurolink.ts (1)
src/lib/types/conversation.ts (1)
ProviderDetails(397-400)
🪛 dotenv-linter (4.0.0)
.env.example
[warning] 325-325: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 406-406: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 407-407: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 408-408: [UnorderedKey] The NEUROLINK_SUMMARIZATION_PROVIDER key should go before the NEUROLINK_TOKEN_THRESHOLD key
(UnorderedKey)
[warning] 408-408: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 409-409: [UnorderedKey] The NEUROLINK_SUMMARIZATION_MODEL key should go before the NEUROLINK_SUMMARIZATION_PROVIDER key
(UnorderedKey)
[warning] 409-409: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
🔇 Additional comments (14)
src/lib/types/streamTypes.ts (1)
220-221: LGTM!The
enableSummarizationproperty addition is well-placed and correctly typed as an optional boolean, enabling per-request control over summarization behavior in streaming operations.src/lib/types/generateTypes.ts (1)
228-229: LGTM!The
enableSummarizationproperty is correctly added toTextGenerationOptions, maintaining consistency with the same property inStreamOptionsfor unified per-request summarization control.src/lib/config/conversationMemory.ts (2)
36-52: LGTM!The token-based memory constants are well-documented and sensibly configured:
- 80% threshold leaves headroom for new messages
- 50k fallback is reasonable for most models
- 30% recent messages ratio balances context continuity with summarization efficiency
59-81: LGTM!The
getConversationMemoryDefaultsfunction properly integrates token-based configuration while maintaining backward compatibility with deprecated turn-based fields. The conditionaltokenThresholdassignment correctly allows dynamic calculation when the environment variable isn't set..env.example (1)
332-343: LGTM - Redis configuration additions.The new Redis configuration variables are comprehensive and follow standard conventions. The documented defaults (TTL of 24 hours, reasonable timeouts and retry settings) are appropriate for production use.
src/lib/utils/conversationMemoryUtils.ts (1)
92-109: LGTM - Clean refactor to options-based API.The transition to an options object for
storeConversationTurnimproves readability and maintainability. TheProviderDetailsconstruction correctly guards against missing provider/model values.src/lib/types/conversation.ts (3)
278-311: KeepConversationBaseandSessionMemoryin sync for token-based fieldsYou’ve added token-based fields (
summarizedUpToMessageId,summarizedMessage,tokenThreshold,lastTokenCount,lastCountedAt) to bothSessionMemoryandConversationBase, which is good for alignment between in-memory and Redis.Make sure all code that mutates these fields updates both representations consistently (e.g., when a Redis object is deserialized into a
SessionMemorywrapper and then written back), as inconsistencies here will directly affectbuildContextFromPointerand token-counting behavior. The current PR mostly does this, but it’s an easy place for future drift—worth keeping an eye on in follow-up changes.
397-400:ProviderDetailstype is clear and minimalThe standalone
ProviderDetailstype (provider + model) is a good fit for both memory and logging metadata, and keeps token-threshold utilities decoupled from broader provider configs. No issues here.
113-152: ChatMessage.id requirement is already properly enforced across the codebaseThe
ChatMessagetype correctly requiresid: string(line 115, conversation.ts), and all actualChatMessageconstructions throughout the codebase properly supply it viarandomUUID()or synthetic IDs:
src/lib/core/redisConversationMemoryManager.ts: User, assistant, tool call, and tool result messages all includeidsrc/lib/utils/conversationMemory.ts: Summary messages use synthetic IDs (e.g.,summary-${sessionId})src/lib/core/conversationMemoryManager.ts:validateAndPrepareMessage()method generates UUIDs for all messagesThe message objects found in other locations (mem0.add calls, provider adapters, CLI factories) are not
ChatMessagetype—they are provider-specific or external library formats. MessageBuilder correctly createsCoreMessage[]arrays for AI SDK compatibility, notChatMessage[]. No modifications needed; the type system already preventsChatMessageconstruction withoutid.src/lib/core/redisConversationMemoryManager.ts (2)
510-651: Token-based summarization pipeline is sound and filters out tool messages appropriately
checkAndSummarizebuilds aSessionMemorywrapper, usesbuildContextFromPointerto respect any existing summary pointer, estimates tokens viaTokenUtils.estimateTokenCount, and triggerssummarizeSessionTokenBasedwhen over threshold.summarizeSessionTokenBasedthen:
- Starts from
summarizedUpToMessageId+ 1 (or 0),- Filters out
tool_call/tool_resultmessages before summarizing,- Uses a token-based split so the most recent ~30% tokens stay as raw messages,
- Updates
summarizedUpToMessageIdandsummarizedMessageand writes back to Redis.This is a solid, non‑destructive design and matches the intended token‑based summarization behavior.
1030-1045:createSummarySystemMessagemetadata aligns with ChatMessage additionsThe updated
createSummarySystemMessagenow sets an explicitid,timestamp, andmetadata.isSummary/summarizesFrom/summarizesTo. This is consistent with the newChatMessageshape and will play nicely with the token-based pointer logic. No changes needed here.src/lib/core/conversationMemoryManager.ts (3)
179-206: Token-based summarization trigger is well-structured and logs useful diagnostics
checkAndSummarize:
- Builds context via
buildContextFromPointer(session),- Updates
session.lastTokenCount/lastCountedAt,- Logs the current count vs threshold,
- Invokes
summarizeSessionTokenBasedwhen appropriate, all within a guardedtry/catch.This is a clean trigger path and provides enough logging to debug threshold behavior. No changes needed.
231-247: Summary system message metadata matches type expectations
createSummarySystemMessagenow creates a summaryChatMessagewith an ID, timestamp, andmetadatacarryingisSummaryand range bounds. This lines up with the newChatMessage.metadatafields and will work correctly with pointer-based context building. Looks good.
309-329:findSplitIndexByTokenslogic is correct and matches the Redis strategyThe function walks messages from the end, accumulating recent tokens until it would exceed
targetRecentTokens, then chooses a split point so at least one message is always summarized. This matches the “keep the most recent ~RECENT_MESSAGES_RATIO tokens as raw messages” design, and looks correct.
14740db to
bc7cf1b
Compare
|
Full review |
|
@BoraYaswanthReddy I'll conduct a full review of this PR, analyzing the token-based summarization implementation across all changed files. ✅ Actions performedFull review triggered. |
60c87c5 to
5cb3e31
Compare
|
@coderabbitai Please review this pr |
|
@BoraYaswanthReddy I'll provide a comprehensive review of this token-based summarization implementation. 🎯 Overall Architecture AssessmentThis PR represents a significant architectural shift from turn-based to token-based conversation memory management. The approach is well-structured with proper separation of concerns. Strengths:
|
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
src/lib/neurolink.ts (1)
3270-3293: Fallback streaming never stores to conversation memory due toenhancedOptionsguardIn
handleStreamError, the fallback stream’sfinallyblock gates memory writes onenhancedOptions?.context?.sessionId, but the only call site passesundefinedforenhancedOptions. As a result, fallback turns are never persisted or summarized, even whenoptions.context.sessionIdis present.You can both pass the real
enhancedOptionsfromstream()and fall back tooptions.contextwhen it’s unavailable:Proposed fix for fallback streaming memory write
--- a/src/lib/neurolink.ts +++ b/src/lib/neurolink.ts @@ -3006,12 +3006,16 @@ async stream(options: StreamOptions): Promise<StreamResult> { } catch (error) { - return this.handleStreamError( - error, - options, - startTime, - streamId, - undefined, - undefined, - ); + return this.handleStreamError( + error, + options, + startTime, + streamId, + enhancedOptions, + factoryResult, + ); } @@ -3268,9 +3272,15 @@ const fallbackProcessedStream = (async function* (self: NeuroLink) { } finally { // Store memory after fallback stream consumption is complete - if (self.conversationMemory && enhancedOptions?.context?.sessionId) { - const sessionId = ( - enhancedOptions?.context as Record<string, unknown> - )?.sessionId as string; - const userId = (enhancedOptions?.context as Record<string, unknown>) - ?.userId as string; + const context = + (enhancedOptions?.context as Record<string, unknown> | undefined) ?? + (options.context as Record<string, unknown> | undefined); + + if (self.conversationMemory && context?.sessionId) { + const sessionId = context.sessionId as string; + const userId = context.userId as string | undefined; @@ -3283,12 +3293,12 @@ const fallbackProcessedStream = (async function* (self: NeuroLink) { try { await self.conversationMemory.storeConversationTurn({ - sessionId: sessionId || (options.context?.sessionId as string), - userId: userId || (options.context?.userId as string), + sessionId, + userId, userMessage: originalPrompt ?? "", aiResponse: fallbackAccumulatedContent, startTimeStamp: new Date(startTime), providerDetails, - enableSummarization: enhancedOptions?.enableSummarization, + enableSummarization: + enhancedOptions?.enableSummarization ?? + options.enableSummarization, });src/lib/core/redisConversationMemoryManager.ts (1)
334-353: In-process summarization guard is good, but multi-node Redis setups still have a race windowThe new
summarizationInProgressset combined withsetImmediatetriggers instoreConversationTurnand thecheckAndSummarizeguard prevents overlapping summarizations per session within a single Node process. However, in a typical Redis-backed deployment with multiple Neurolink instances, each process has its ownsummarizationInProgressset; concurrent requests hitting different nodes can still:
- Trigger multiple summarizations for the same
{sessionId,userId}in parallel, and- Race writing updated
RedisConversationObjects back to Redis (last writer wins onsummarizedUpToMessageId/summarizedMessage/token counts).Functionally this “just” causes redundant summarization and potentially non-monotonic pointer updates, but given Redis is intended for distributed memory, it’s worth addressing.
You may want to introduce a lightweight Redis-based lock (e.g.,
SETNXwith TTL on asummary-lock:${sessionId}:${userId}key) aroundcheckAndSummarize/summarizeSessionTokenBasedso only one node can summarize a session at a time, while keeping the in-processsummarizationInProgressas a fast local guard.Also applies to: 429-437, 466-502, 538-598
🧹 Nitpick comments (9)
src/lib/config/conversationMemory.ts (1)
88-95: Consider adding@deprecatedJSDoc tags for IDE visibility.The inline comment
// Deprecated (for backward compatibility)won't surface in IDE tooltips or generate warnings. Adding proper@deprecatedJSDoc tags would improve developer experience:+ /** @deprecated Use tokenThreshold instead */ maxTurnsPerSession: Number(process.env.NEUROLINK_MEMORY_MAX_TURNS_PER_SESSION) || DEFAULT_MAX_TURNS_PER_SESSION, + /** @deprecated Use tokenThreshold instead */ summarizationThresholdTurns: Number(process.env.NEUROLINK_SUMMARIZATION_THRESHOLD_TURNS) || 20, + /** @deprecated Use tokenThreshold instead */ summarizationTargetTurns: Number(process.env.NEUROLINK_SUMMARIZATION_TARGET_TURNS) || 10,src/lib/utils/conversationMemory.ts (3)
231-248: Handle edge case where pointer message may be stale or removed.The implementation correctly falls back to all messages when the pointer is not found. However, this could mask data integrity issues in long-running sessions.
Consider logging at a higher severity or tracking this as a metric since a missing pointer suggests the session state may be inconsistent (e.g., messages were deleted but pointer wasn't cleared).
348-376: Consider adding upper-bound validation for threshold overrides.The function validates that overrides are positive (
> 0) but doesn't cap unreasonably high values. An override exceeding the model's actual context window could lead to unexpected behavior.Suggested validation
export function getEffectiveTokenThreshold( provider: string, model: string, envOverride?: number, sessionOverride?: number, ): number { + const modelLimit = calculateTokenThreshold(provider, model); + const maxAllowed = modelLimit * 1.5; // Allow some headroom but cap extreme values + // Priority 1: Session-level override if (sessionOverride && sessionOverride > 0) { + if (sessionOverride > maxAllowed) { + logger.warn("Session threshold override exceeds model limit, capping", { + requested: sessionOverride, + capped: maxAllowed, + }); + return maxAllowed; + } return sessionOverride; }
397-399: New NeuroLink instance created per summarization call.Creating a new
NeuroLinkinstance for each summary generation works but may have overhead for frequent summarization. If performance becomes a concern, consider reusing a singleton summarizer instance or caching the NeuroLink instance at the module level.src/lib/types/conversation.ts (1)
21-47: Config-level move to token thresholds with deprecated turn-based knobs looks goodAdding
tokenThreshold?: numberand marking the older turn-based knobs as@deprecatedmatches the new token-based summarization design while preserving backward compatibility at the type level. Consider also updating higher-level docs/usages (e.g., constructor JSDoc inNeuroLink) to steer callers towardtokenThresholdover the legacy turn-based fields.src/lib/core/conversationMemoryManager.ts (1)
150-190: Minor behavioral differences vs Redis manager (tool messages, ratio constant)The in‑memory manager’s
summarizeSessionTokenBasedcurrently:
- Includes all message roles (user/assistant/system/tool_call/tool_result) in
recentMessages, and- Uses
RECENT_MESSAGES_RATIOfrom config for the recent‑token budget,while the Redis manager filters out tool_call/tool_result messages and hardcodes
0.3as the ratio. Behavior is correct here, but for predictability it would be cleaner to:
- Reuse the same ratio constant in both managers, and
- Decide consistently whether tool_* messages should be part of the summarized region.
Also applies to: 192-238, 284-336, 341-361
src/lib/core/redisConversationMemoryManager.ts (2)
612-673: Unify summarization behavior and constants with in-memory managerIn
summarizeSessionTokenBasedyou:
- Filter
recentMessagesto excludetool_callandtool_resultroles, and- Use a hardcoded
threshold * 0.3fortargetRecentTokens,while the in-memory
ConversationMemoryManager:
- Summarizes all roles, and
- Uses
RECENT_MESSAGES_RATIOfromconfig/conversationMemory.ts.The current behavior is not wrong, but it’s asymmetrical and duplicates the ratio. To keep both backends aligned and easier to tune, consider:
- Importing and using
RECENT_MESSAGES_RATIOhere instead of the literal0.3, and- Deciding consistently (both managers) whether tool_* messages should be summarized or always kept in the recent window.
Also applies to: 678-698
705-756: Redis context building and message IDs look consistent with new typesThe updated
buildContextMessagescorrectly:
- Rehydrates a
SessionMemoryfrom the stored RedisConversationObject,- Uses
buildContextFromPointerto inject summaries, and- Optionally filters tool_call/tool_result when summarization is enabled.
Similarly,
createSummarySystemMessageandflushPendingToolDatanow assign UUIDids and appropriate metadata to summary/tool messages, which matches the newChatMessagecontract and pointer-based summarization logic.Also applies to: 1077-1093, 1359-1452
src/lib/neurolink.ts (1)
110-116: Top-levelawaitis safely supported in Node 20+ ESM; moving to lazy initializer is optionalTop-level await is natively available in Node 14.8+ when using ES modules, and your project targets Node 20.9.0+ with
"type": "module", so this pattern incurs no compatibility constraints. If you prefer deferring initialization, moving the dynamic import intolazyInitializeConversationMemoryis a valid stylistic choice, but not necessary.
📜 Review details
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (15)
.env.example(2 hunks)src/cli/loop/optionsSchema.ts(1 hunks)src/lib/config/conversationMemory.ts(2 hunks)src/lib/core/conversationMemoryFactory.ts(0 hunks)src/lib/core/conversationMemoryInitializer.ts(0 hunks)src/lib/core/conversationMemoryManager.ts(2 hunks)src/lib/core/redisConversationMemoryManager.ts(10 hunks)src/lib/neurolink.ts(5 hunks)src/lib/types/conversation.ts(8 hunks)src/lib/types/generateTypes.ts(1 hunks)src/lib/types/sdkTypes.ts(0 hunks)src/lib/types/streamTypes.ts(1 hunks)src/lib/utils/conversationMemory.ts(4 hunks)src/lib/utils/conversationMemoryUtils.ts(2 hunks)src/lib/utils/redis.ts(0 hunks)
💤 Files with no reviewable changes (4)
- src/lib/types/sdkTypes.ts
- src/lib/utils/redis.ts
- src/lib/core/conversationMemoryFactory.ts
- src/lib/core/conversationMemoryInitializer.ts
🚧 Files skipped from review as they are similar to previous changes (2)
- src/lib/types/streamTypes.ts
- src/lib/types/generateTypes.ts
🧰 Additional context used
📓 Path-based instructions (2)
**/*.{ts,tsx}
📄 CodeRabbit inference engine (CLAUDE.md)
**/*.{ts,tsx}: Maintain strict TypeScript type safety across all modules with comprehensive type definitions organized by domain to avoid circular dependencies
Use ErrorFactory for creating typed errors throughout the application
Wrap async operations with withTimeout utility for timeout handling
Files:
src/cli/loop/optionsSchema.tssrc/lib/utils/conversationMemory.tssrc/lib/utils/conversationMemoryUtils.tssrc/lib/core/conversationMemoryManager.tssrc/lib/config/conversationMemory.tssrc/lib/core/redisConversationMemoryManager.tssrc/lib/neurolink.tssrc/lib/types/conversation.ts
**/types/**/*.ts
📄 CodeRabbit inference engine (CLAUDE.md)
Type definitions must be organized by domain (providers, generation, streaming, MCP, etc.) to avoid circular dependencies
Files:
src/lib/types/conversation.ts
🧠 Learnings (14)
📓 Common learnings
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Memory management should use Redis for distributed memory in production and in-memory store for development, with conversation summarization for long contexts
Learnt from: BoraYaswanthReddy
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-12-14T18:33:26.766Z
Learning: In the neurolink repository's Redis conversation memory implementation, message IDs were changed from sequential integers to UUIDs (using `generateUniqueId()`) as a security improvement to prevent enumeration attacks and information leakage about conversation volumes and patterns.
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Memory management should use Redis for distributed memory in production and in-memory store for development, with conversation summarization for long contexts
Applied to files:
src/lib/utils/conversationMemory.tssrc/lib/core/conversationMemoryManager.tssrc/lib/config/conversationMemory.tssrc/lib/core/redisConversationMemoryManager.ts.env.examplesrc/lib/types/conversation.ts
📚 Learning: 2025-09-24T06:42:06.088Z
Learnt from: amreetkhuntia
Repo: juspay/neurolink PR: 185
File: src/lib/evaluation/contextBuilder.ts:79-85
Timestamp: 2025-09-24T06:42:06.088Z
Learning: In the NeuroLink codebase, using `(options.prompt || [])` pattern for handling potentially undefined prompt arrays is the preferred approach over extracting to a normalized variable when building conversation history in the ContextBuilder class.
Applied to files:
src/lib/utils/conversationMemory.tssrc/lib/neurolink.ts
📚 Learning: 2025-09-24T07:26:41.988Z
Learnt from: amreetkhuntia
Repo: juspay/neurolink PR: 185
File: src/lib/evaluation/prompts.ts:86-101
Timestamp: 2025-09-24T07:26:41.988Z
Learning: In the neurolink codebase, maintainer amreetkhuntia consistently prefers to keep template literal indentation in LLM prompts (including evaluation prompts in src/lib/evaluation/prompts.ts) for readability, even when it results in extra whitespace in the output, as LLMs can parse and understand the content correctly.
Applied to files:
src/lib/utils/conversationMemory.ts.env.example
📚 Learning: 2025-09-24T06:43:23.653Z
Learnt from: amreetkhuntia
Repo: juspay/neurolink PR: 185
File: src/lib/evaluation/prompts.ts:59-72
Timestamp: 2025-09-24T06:43:23.653Z
Learning: In the neurolink codebase, maintainer amreetkhuntia prefers to keep template literal indentation in LLM prompts even if it results in technically malformed JSON format, as LLMs can understand and parse it correctly despite formatting issues.
Applied to files:
src/lib/utils/conversationMemory.ts.env.example
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/cli/loop/session.ts : Loop mode interactive sessions should be implemented in src/cli/loop/session.ts with persistent conversation memory and session-wide configuration
Applied to files:
src/lib/utils/conversationMemory.tssrc/lib/core/conversationMemoryManager.tssrc/lib/core/redisConversationMemoryManager.tssrc/lib/types/conversation.ts
📚 Learning: 2025-09-17T17:55:15.261Z
Learnt from: RajuSudhar
Repo: juspay/neurolink PR: 173
File: src/lib/index.ts:16-16
Timestamp: 2025-09-17T17:55:15.261Z
Learning: In src/lib/types/providers.ts, ProviderConfig was renamed to AIModelProviderConfig to deduplicate type names, as there was an existing ProviderConfig type that better suited the "ProviderConfig" name. This was an intentional breaking change for better type organization.
Applied to files:
src/lib/utils/conversationMemoryUtils.tssrc/lib/neurolink.tssrc/lib/types/conversation.ts
📚 Learning: 2025-12-14T18:33:26.766Z
Learnt from: BoraYaswanthReddy
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-12-14T18:33:26.766Z
Learning: In the neurolink repository's Redis conversation memory implementation, message IDs were changed from sequential integers to UUIDs (using `generateUniqueId()`) as a security improvement to prevent enumeration attacks and information leakage about conversation volumes and patterns.
Applied to files:
src/lib/core/redisConversationMemoryManager.ts.env.example
📚 Learning: 2025-12-12T20:11:17.070Z
Learnt from: Yaswanth-2874
Repo: juspay/neurolink PR: 672
File: src/lib/core/redisConversationMemoryManager.ts:1082-1091
Timestamp: 2025-12-12T20:11:17.070Z
Learning: In the Redis conversation memory implementation, LLM context keys (when `separateLLMContext` is enabled) intentionally use only sessionId without userId: `llm:context:${sessionId}`. This is by design to scope LLM context purely at the session level, relying on sessionId global uniqueness.
Applied to files:
src/lib/core/redisConversationMemoryManager.ts
📚 Learning: 2025-09-17T18:14:34.960Z
Learnt from: RajuSudhar
Repo: juspay/neurolink PR: 173
File: src/lib/types/index.ts:58-62
Timestamp: 2025-09-17T18:14:34.960Z
Learning: RajuSudhar explained that in the Neurolink codebase, there are multiple ProviderConfig types causing inconsistency. One existing ProviderConfig type better suited the "ProviderConfig" name, so they renamed the less-suitable one to AIModelProviderConfig to free up the name. Adding backward compatibility aliases would worsen naming inconsistency rather than help. The remaining duplicates will be systematically deduplicated in the 07-Types-Module.md TODO as part of their phased refactor approach.
Applied to files:
src/lib/neurolink.ts
📚 Learning: 2025-12-14T18:33:26.766Z
Learnt from: BoraYaswanthReddy
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-12-14T18:33:26.766Z
Learning: In the neurolink repository, the `separateLLMContext` configuration exists primarily to support agentic loops that need full conversation context including tool messages. The default is intentionally `true` because separating tool messages from LLM context is considered the better default behavior. For CLI usage, separation is always enabled and the option is not exposed to users.
Applied to files:
.env.example
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/utils/messageBuilder.ts : Message construction must be handled through MessageBuilder in src/lib/utils/messageBuilder.ts, which handles text, images, PDFs, and CSV files with provider-specific adapters
Applied to files:
src/lib/types/conversation.ts
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/providers/*.ts : Providers must extend a base provider or implement the provider interface and register in ProviderRegistry.registerAllProviders() with provider name, factory function, default model, and aliases
Applied to files:
src/lib/types/conversation.ts
📚 Learning: 2025-12-10T12:24:51.147Z
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-12-10T12:24:51.147Z
Learning: Applies to **/types/index.ts : Add new provider names to the AIProviderName enum in src/lib/types/index.ts when adding a new provider
Applied to files:
src/lib/types/conversation.ts
🧬 Code graph analysis (4)
src/lib/utils/conversationMemory.ts (3)
src/lib/types/conversation.ts (3)
ProviderDetails(381-384)SessionMemory(52-97)ChatMessage(113-152)src/lib/config/conversationMemory.ts (2)
MEMORY_THRESHOLD_PERCENTAGE(40-40)DEFAULT_FALLBACK_THRESHOLD(45-45)src/lib/neurolink.ts (1)
NeuroLink(152-6117)
src/lib/utils/conversationMemoryUtils.ts (1)
src/lib/types/conversation.ts (1)
ProviderDetails(381-384)
src/lib/core/conversationMemoryManager.ts (3)
src/lib/types/conversation.ts (4)
SessionMemory(52-97)ConversationMemoryConfig(11-47)StoreConversationTurnOptions(237-245)ChatMessage(113-152)src/lib/utils/conversationMemory.ts (3)
getEffectiveTokenThreshold(348-376)buildContextFromPointer(231-273)generateSummary(387-419)src/lib/config/conversationMemory.ts (2)
MEMORY_THRESHOLD_PERCENTAGE(40-40)RECENT_MESSAGES_RATIO(52-52)
src/lib/neurolink.ts (1)
src/lib/types/conversation.ts (1)
ProviderDetails(381-384)
🪛 dotenv-linter (4.0.0)
.env.example
[warning] 325-325: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 406-406: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 407-407: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 408-408: [UnorderedKey] The NEUROLINK_SUMMARIZATION_PROVIDER key should go before the NEUROLINK_TOKEN_THRESHOLD key
(UnorderedKey)
[warning] 408-408: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
[warning] 409-409: [UnorderedKey] The NEUROLINK_SUMMARIZATION_MODEL key should go before the NEUROLINK_SUMMARIZATION_PROVIDER key
(UnorderedKey)
[warning] 409-409: [ValueWithoutQuotes] This value needs to be surrounded in quotes
(ValueWithoutQuotes)
🔇 Additional comments (13)
src/cli/loop/optionsSchema.ts (1)
90-94: LGTM!The new
enableSummarizationoption follows the established pattern for boolean CLI options and aligns with per-request summarization control introduced across the codebase.src/lib/utils/conversationMemoryUtils.ts (1)
92-110: Well-structured options object pattern.The refactoring to use a single options object for
storeConversationTurnimproves API clarity and extensibility. The conditional construction ofproviderDetailscorrectly guards against missing provider/model values..env.example (2)
324-325: Note: Default behavior change for memory feature.
NEUROLINK_MEMORY_ENABLEDis nowtrueby default. This changes the out-of-box behavior for new installations. Ensure this is documented in release notes as users who previously relied on memory being disabled by default may experience unexpected behavior.
404-413: Token-based summarization configuration looks good.The documentation clearly explains the token-threshold approach with the 80% model context default. The deprecated turn-based settings are properly marked and commented out for backward compatibility reference.
src/lib/config/conversationMemory.ts (2)
36-52: Well-documented constants for token-based memory.The new constants are clearly documented and provide sensible defaults. The 80% threshold and 30% recent messages ratio align with the PR objectives for provider-aware threshold calculation.
78-79: Consider the default-true behavior forenableSummarization.The condition
!== "false"means summarization is enabled by default when the env var is unset or set to any value other than"false". This is intentional but differs from the pattern used forenabled(which requires=== "true"). Verify this asymmetry is desired.src/lib/utils/conversationMemory.ts (3)
253-262: Good pointer-based context construction.The summary message correctly uses role
"system"and includes proper metadata (isSummary,summarizesTo) for downstream processing. The generated ID patternsummary-${pointerId}is deterministic and traceable.
78-83: Clean integration of per-request summarization toggle.The
enableSummarizationis correctly extracted with nullish coalescing and passed through tobuildContextMessages, enabling per-request control over summarization behavior.
184-202: Well-implemented conversation turn storage with provider metadata.The code correctly:
- Normalizes
aiResponsewith nullish coalescing (?? "")- Conditionally constructs
providerDetailsonly when both provider and model are available- Passes the complete options object to
storeConversationTurnThis aligns with the
StoreConversationTurnOptionstype and the broader refactoring pattern in this PR.src/lib/neurolink.ts (1)
2931-2954: Streaming path memory write wiring looks correctThe main streaming path now stores conversation turns via
conversationMemory.storeConversationTurnwithsessionId,userId,userMessage(original prompt),aiResponse(accumulated stream),startTimeStamp,providerDetails, andenableSummarization. This matchesStoreConversationTurnOptionsand correctly forwards per‑requestenableSummarizationand provider/model metadata to the memory layer.src/lib/types/conversation.ts (2)
234-245:StoreConversationTurnOptionsandProviderDetailsare well-shaped for public APIThe new
StoreConversationTurnOptionscleanly captures the data needed by both in‑memory and Redis managers, andproviderDetails/enableSummarizationgive the memory layer enough context for provider‑aware thresholds and per‑request overrides. The standaloneProviderDetailstype keeps this reusable and is already being consumed correctly fromNeuroLink.Also applies to: 381-384
113-152: AllChatMessageinstantiations already include requiredidfield—no action neededVerification confirms that every ChatMessage object created across conversationMemoryManager.ts, redisConversationMemoryManager.ts, and conversationMemory.ts properly sets the
idfield. User and assistant messages userandomUUID(), while summary messages usesummary-${uuid}format. The initial concern was based on false positives from external Mem0 library calls and other message types (CoreMessage, MultimodalMessage) that are distinct from ChatMessage.Likely an incorrect or invalid review comment.
src/lib/core/conversationMemoryManager.ts (1)
66-144: Token-basedstoreConversationTurnpath and async summarization trigger look solidSwitching
storeConversationTurntoStoreConversationTurnOptions, computing a provider‑awaretokenThresholdviagetEffectiveTokenThreshold, validating/truncating messages withTokenUtils, and kicking off background summarization throughcheckAndSummarize(guarded bysummarizationInProgress) gives the in‑memory manager a coherent token‑based flow and avoids the previous empty‑array summarization pitfall.
|
Even though id was optional earlier, we still generated id for every conversation Message |
|
Tip For best results, initiate chat on the files or code changes. @BoraYaswanthReddy, thank you for those clarifications! That context is very helpful. On the On token estimation accuracy: You're absolutely right about the speed vs. accuracy trade-off. Reducing latency can significantly enhance user experience, but it often comes at the cost of model quality, and since token estimation happens after every message, intelligent context management requires controlling what information reaches the LLM while preserving the signal your application needs for accurate responses. Your character-based approximation (chars/4) provides fast estimation with acceptable accuracy for the use case, which is the right engineering decision for frequent operations. Both of these are excellent clarifications that address my earlier concerns. The remaining items from my review (race conditions, error handling, bounds checking, etc.) are still worth considering, but these two points are well-justified. 🐰✨ ✏️ Learnings added
🧠 Learnings used |
- Add support for token based summarization - Add param for generate and stream to enable/disable token based summarization
5cb3e31 to
a81cf82
Compare
Pull Request
Description
This PR implements token-based automatic conversation summarization for NeuroLink's conversation memory system. The implementation intelligently manages long conversations by automatically summarizing older messages when token limits are approached, while keeping recent messages intact for context continuity.
Type of Change
Related Issues
Changes Made
Core Features
enableSummarizationoption to bothStreamOptionsandTextGenerationOptionsfor fine-grained controlAPI Enhancements
enableSummarization?: booleanfieldenableSummarization?: booleanfieldenableSummarizationoverrides instance-level configurationImplementation Details
storeConversationTurn()to use options object pattern (was exceeding 6-parameter limit)ConversationMemoryManager(in-memory) andRedisConversationMemoryManager(Redis)conversationMemory.ts,conversationMemoryUtils.ts) to support new signaturesetImmediate()to avoid blocking main operationsConfiguration Changes
enableSummarizationstatusAI Provider Impact
Note: Summarization feature works with all AI providers that support text generation. The summarization uses a configurable provider/model (defaults to the same provider being used).
Component Impact
Testing
Manual Testing Performed
Test Scenarios Validated
enableSummarization: false→ skips summarization even if instance config enables itenableSummarization: true→ enables summarization even if instance config disables itTest Environment
Performance Impact
Benefits
Overhead
Breaking Changes
None - This is a backward-compatible addition. All changes are opt-in:
enableSummarizationcontinues to work unchangedChecklist
Additional Notes
Priority System
The summarization control follows this priority order:
enableSummarization(highest priority)conversationMemory.enableSummarizationconfigfalse(no summarization)Architecture Highlights
StoreConversationTurnOptionsobject to comply with linting rulesFuture Enhancements
Summary by CodeRabbit
Release Notes
New Features
Configuration Changes
✏️ Tip: You can customize this high-level summary in your review settings.