Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -369,6 +369,29 @@ When changing how messages are constructed:
4. Add tests for new message types
5. Update type definitions in `src/lib/types/conversation.ts`

### Embeddings

Providers expose `embed()` and `embedMany()` methods for generating vector embeddings. The `AIProvider` type interface includes both methods; unsupported providers throw descriptive errors.

**Supported providers and defaults:**

| Provider | Default Model | Env Override |
| ---------------- | ------------------------------ | --------------------------- |
| OpenAI | `text-embedding-3-small` | — |
| Google AI Studio | `gemini-embedding-001` | `GOOGLE_AI_EMBEDDING_MODEL` |
| Google Vertex | `text-embedding-004` | `VERTEX_EMBEDDING_MODEL` |
| Amazon Bedrock | `amazon.titan-embed-text-v2:0` | — |

**Server routes:** `POST /api/agent/embed` (single) and `POST /api/agent/embed-many` (batch) in `src/lib/server/routes/agentRoutes.ts`.

**Key files:**

- `src/lib/core/baseProvider.ts` — Base `embed()` / `embedMany()` stubs
- `src/lib/types/providers.ts` — `AIProvider` type with embedding methods
- `src/lib/server/routes/agentRoutes.ts` — Server embedding endpoints
- `src/lib/server/utils/validation.ts` — `EmbedRequestSchema` / `EmbedManyRequestSchema`
- `src/lib/server/types.ts` — `EmbedRequest`, `EmbedResponse`, `EmbedManyRequest`, `EmbedManyResponse`

### Working with Multimodal Content

**For images:**
Expand Down
57 changes: 57 additions & 0 deletions docs/advanced/api-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,63 @@ POST /api/stream
GET /api/status
```

## Embeddings

### Generate Embedding

```http
POST /api/agent/embed
```

**Request body:**

```json
{
"text": "Hello world",
"provider": "googleAiStudio",
"model": "gemini-embedding-001"
}
```

**Response:**

```json
{
"embedding": [0.123, -0.456, ...],
"provider": "googleAiStudio",
"model": "gemini-embedding-001",
"dimension": 768
}
```

### Generate Batch Embeddings

```http
POST /api/agent/embed-many
```

**Request body:**

```json
{
"texts": ["First document", "Second document", "Third document"],
"provider": "openai",
"model": "text-embedding-3-small"
}
```

**Response:**

```json
{
"embeddings": [[0.123, -0.456, ...], [0.789, -0.012, ...], [0.345, -0.678, ...]],
"provider": "openai",
"model": "text-embedding-3-small",
"count": 3,
"dimension": 1536
}
```
Comment on lines +25 to +80

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

This page now mixes two different API surfaces.

The new section documents /api/agent/* endpoints, but the same page still lists /api/generate, /api/stream, and /api/status above. That makes the reference internally inconsistent and likely sends readers to routes that do not exist on the current server API. Please either migrate the older sections to the same namespace or split legacy/current APIs explicitly.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@docs/advanced/api-reference.md` around lines 25 - 80, The docs page mixes two
API surfaces: the new /api/agent/* endpoints (e.g., /api/agent/embed,
/api/agent/embed-many) and older endpoints (/api/generate, /api/stream,
/api/status); fix by choosing one of two approaches—either migrate the legacy
endpoints to the /api/agent namespace (update all examples, request/response
fields, and references to use /api/agent/generate, /api/agent/stream,
/api/agent/status) or split the page into clearly labeled sections "Current API
(/api/agent/...)" and "Legacy API (/api/...)" with a short deprecation note and
consistent examples for each section—ensure all endpoint paths and example
bodies reference the chosen namespace consistently.


## MCP Integration

### List MCP Tools
Expand Down
44 changes: 44 additions & 0 deletions docs/api/type-aliases/AIProvider.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,50 @@ Defined in: [types/providers.ts:308](https://github.com/juspay/neurolink/blob/1b

---

### embed()

> **embed**(`text`, `modelName?`): `Promise`\<`number[]`\>

Generate an embedding vector for a single text. Throws if the provider does not support embeddings.

#### Parameters

##### text

`string`

##### modelName?

`string`

#### Returns

`Promise`\<`number[]`\>

---

### embedMany()

> **embedMany**(`texts`, `modelName?`): `Promise`\<`number[][]`\>

Generate embedding vectors for multiple texts in a single batch. The AI SDK automatically handles chunking for models with batch limits.

#### Parameters

##### texts

`string[]`

##### modelName?

`string`

#### Returns

`Promise`\<`number[][]`\>

---

### setupToolExecutor()

> **setupToolExecutor**(`sdk`, `functionTag`): `void`
Expand Down
65 changes: 60 additions & 5 deletions docs/guides/server-adapters/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,11 +147,13 @@ All server adapters expose the same REST API endpoints:

### Agent Operations

| Endpoint | Method | Description |
| ---------------------- | ------ | -------------------------------------- |
| `/api/agent/execute` | POST | Execute agent and return full response |
| `/api/agent/stream` | POST | Stream agent response via SSE |
| `/api/agent/providers` | GET | List available AI providers |
| Endpoint | Method | Description |
| ----------------------- | ------ | -------------------------------------------- |
| `/api/agent/execute` | POST | Execute agent and return full response |
| `/api/agent/stream` | POST | Stream agent response via SSE |
| `/api/agent/providers` | GET | List available AI providers |
| `/api/agent/embed` | POST | Generate embedding for a single text |
| `/api/agent/embed-many` | POST | Generate embeddings for multiple texts batch |

### Tool Operations

Expand Down Expand Up @@ -433,6 +435,59 @@ data: {"type":"text-end","timestamp":1706745600100}
data: {"type":"finish","usage":{"inputTokens":5,"outputTokens":50,"totalTokens":55}}
```

### Generate Embedding

**Request:**

```json
POST /api/agent/embed
Content-Type: application/json

{
"text": "What is the meaning of life?",
"provider": "openai",
"model": "text-embedding-3-small"
}
```

**Response:**

```json
{
"embedding": [0.123, -0.456, 0.789, ...],
"provider": "openai",
"model": "text-embedding-3-small",
"dimension": 1536
}
```

### Generate Batch Embeddings

**Request:**

```json
POST /api/agent/embed-many
Content-Type: application/json

{
"texts": ["First document", "Second document"],
"provider": "googleAiStudio",
"model": "gemini-embedding-001"
}
```

**Response:**

```json
{
"embeddings": [[0.123, -0.456, ...], [0.789, -0.012, ...]],
"provider": "googleAiStudio",
"model": "gemini-embedding-001",
"count": 2,
"dimension": 768
}
```

---

## Production Deployment
Expand Down
41 changes: 41 additions & 0 deletions docs/sdk/api-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -454,6 +454,47 @@ const result = await neurolink.gen({

---

### Embeddings

Generate embeddings directly via the provider's `embed()` and `embedMany()` methods.

#### `provider.embed(text, modelName?)`

Generate an embedding vector for a single text.

```typescript
import { ProviderFactory } from "@juspay/neurolink";

const provider = await ProviderFactory.createProvider("googleAiStudio");
const embedding = await provider.embed("Hello world");
// embedding: number[] (e.g., 768 dimensions)
```

#### `provider.embedMany(texts, modelName?)`

Generate embedding vectors for multiple texts in a single batch. The AI SDK automatically handles chunking for models with batch limits.

```typescript
const provider = await ProviderFactory.createProvider("openai");
const embeddings = await provider.embedMany([
"First document",
"Second document",
"Third document",
]);
// embeddings: number[][] (e.g., 3 × 1536 dimensions)
```

**Supported providers and default models:**

| Provider | Default Embedding Model | Env Override |
| ---------------- | ------------------------------ | --------------------------- |
| OpenAI | `text-embedding-3-small` | — |
| Google AI Studio | `gemini-embedding-001` | `GOOGLE_AI_EMBEDDING_MODEL` |
| Google Vertex | `text-embedding-004` | `VERTEX_EMBEDDING_MODEL` |
| Amazon Bedrock | `amazon.titan-embed-text-v2:0` | — |

Comment on lines +487 to +495

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Document these embedding override vars in the main env section.

This table introduces GOOGLE_AI_EMBEDDING_MODEL and VERTEX_EMBEDDING_MODEL, but the later Environment Configuration section in this same file does not list them. Users looking there for the supported env vars will miss these overrides.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@docs/sdk/api-reference.md` around lines 487 - 495, The docs table adds
GOOGLE_AI_EMBEDDING_MODEL and VERTEX_EMBEDDING_MODEL but the Environment
Configuration section doesn't list them; update that main env section to
document these two variables (GOOGLE_AI_EMBEDDING_MODEL,
VERTEX_EMBEDDING_MODEL), include their purpose as embedding model overrides for
Google AI Studio and Google Vertex respectively, indicate default behavior when
unset, and add them to any ENV key/value table or example so readers can
discover and use these overrides.

---

### RAG Integration

Pass `rag: { files: [...] }` to `generate()` or `stream()` for automatic RAG pipeline setup:
Expand Down
27 changes: 27 additions & 0 deletions src/lib/core/baseProvider.ts
Original file line number Diff line number Diff line change
Expand Up @@ -674,14 +674,14 @@
* IMPLEMENTATION NOTE: Uses streamText() under the hood and accumulates results
* for consistency and better performance
*/
async generate(

Check warning on line 677 in src/lib/core/baseProvider.ts

View workflow job for this annotation

GitHub Actions / 🛡️ Code Quality & Security Gate

Async method 'generate' has too many lines (317). Maximum allowed is 300

Check warning on line 677 in src/lib/core/baseProvider.ts

View workflow job for this annotation

GitHub Actions / test (20)

Async method 'generate' has too many lines (317). Maximum allowed is 300
optionsOrPrompt: TextGenerationOptions | string,
_analysisSchema?: ValidationSchema,
): Promise<EnhancedGenerateResult | null> {
return providerTracer.startActiveSpan(
"neurolink.provider.generate",
{ kind: SpanKind.INTERNAL },
async (span) => {

Check warning on line 684 in src/lib/core/baseProvider.ts

View workflow job for this annotation

GitHub Actions / 🛡️ Code Quality & Security Gate

Async arrow function has too many lines (308). Maximum allowed is 300

Check warning on line 684 in src/lib/core/baseProvider.ts

View workflow job for this annotation

GitHub Actions / test (20)

Async arrow function has too many lines (308). Maximum allowed is 300
const options = this.normalizeTextOptions(optionsOrPrompt);
this.validateOptions(options);
const startTime = Date.now();
Expand Down Expand Up @@ -1081,6 +1081,33 @@
);
}

/**
* Generate embeddings for multiple texts in a single batch
*
* This is a default implementation that throws an error.
* Providers that support embeddings should override this method.
* The AI SDK's embedMany automatically handles chunking for models with batch limits.
*
* @param texts - The texts to embed
* @param _modelName - Optional embedding model name (provider-specific)
* @returns Promise resolving to an array of embedding vectors
* @throws Error if the provider does not support embeddings
*/
async embedMany(texts: string[], _modelName?: string): Promise<number[][]> {
logger.warn(
`embedMany() called on ${this.providerName} which does not have a native implementation`,
{
count: texts.length,
},
);
throw new Error(
`Batch embedding generation is not supported by the ${this.providerName} provider. ` +
`Supported providers: openai, googleAiStudio, vertex/google, bedrock. ` +
`Use an embedding model like text-embedding-3-small (OpenAI), gemini-embedding-001 (Google AI), ` +
`text-embedding-004 (Vertex), or amazon.titan-embed-text-v2:0 (Bedrock).`,
);
Comment on lines +1096 to +1108

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Harden embedMany() before exposing it via HTTP.

The new signature has no way to propagate timeout/abort control, and the fallback throws a raw Error when batching is unsupported. That leaves /api/agent/embed-many prone to hung requests on provider stalls and generic 500s for capability misses; please carry timeout/abort through this API and use ErrorFactory for the unsupported-provider path.

As per coding guidelines, "All async operations should be wrapped with withTimeout utility for consistent timeout handling" and "Use ErrorFactory for creating typed errors instead of throwing raw Error objects"

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/core/baseProvider.ts` around lines 1096 - 1108, The embedMany method
on BaseProvider lacks timeout/abort propagation and throws a raw Error for
unsupported batch embedding; wrap the async work in the withTimeout (or
withTimeoutAbortSignal) utility so callers can enforce timeouts/abort signals
and accept an optional AbortSignal/timeout param, and replace the thrown Error
with a typed error created by ErrorFactory (e.g., ErrorFactory.create or the
project-standard factory) that indicates "Batch embedding not supported by
providerName" and includes providerName and supported provider list; update the
embedMany signature to accept and pass through timeout/abort controls and ensure
the logger call remains but the failure path uses the ErrorFactory-created
error.

}

/**
* Get the default embedding model for this provider
*
Expand Down
39 changes: 39 additions & 0 deletions src/lib/providers/amazonBedrock.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2099,4 +2099,43 @@ export class AmazonBedrockProvider extends BaseProvider {
throw this.handleProviderError(error);
}
}

/**
* Generate embeddings for multiple texts in a single batch
* @param texts - The texts to embed
* @param modelName - The embedding model to use (default: amazon.titan-embed-text-v2:0)
* @returns Promise resolving to an array of embedding vectors
*/
async embedMany(texts: string[], modelName?: string): Promise<number[][]> {
const embeddingModelName = modelName || "amazon.titan-embed-text-v2:0";

logger.debug("Generating batch embeddings", {
provider: this.providerName,
model: embeddingModelName,
count: texts.length,
});

try {
const embeddings = await Promise.all(
texts.map((text) => this.embed(text, embeddingModelName)),
);

logger.debug("Batch embeddings generated successfully", {
provider: this.providerName,
model: embeddingModelName,
count: embeddings.length,
embeddingDimension: embeddings[0]?.length,
});

return embeddings;
Comment on lines +2109 to +2130

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Bound batch fan-out before calling Bedrock.

Promise.all sends one Bedrock request per text with no concurrency cap. Large embed-many payloads will spike outbound calls, which makes throttling and timeout storms much more likely here. Also, this duplicates the default model literal instead of reusing getDefaultEmbeddingModel().

♻️ Suggested change
   async embedMany(texts: string[], modelName?: string): Promise<number[][]> {
-    const embeddingModelName = modelName || "amazon.titan-embed-text-v2:0";
+    const embeddingModelName = modelName || this.getDefaultEmbeddingModel();
@@
-      const embeddings = await Promise.all(
-        texts.map((text) => this.embed(text, embeddingModelName)),
-      );
+      const embeddings: number[][] = [];
+      const batchSize = 5;
+
+      for (let i = 0; i < texts.length; i += batchSize) {
+        const batch = texts.slice(i, i + batchSize);
+        embeddings.push(
+          ...(await Promise.all(
+            batch.map((text) => this.embed(text, embeddingModelName)),
+          )),
+        );
+      }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/providers/amazonBedrock.ts` around lines 2109 - 2130, The embedMany
function currently fans out with Promise.all causing unbounded concurrent
Bedrock requests and duplicates the default model literal; change embedMany to
derive the model via getDefaultEmbeddingModel() when modelName is absent, and
replace Promise.all(texts.map(...)) with a controlled concurrent mapper (e.g.,
use a limited-concurrency iterator or p-map-style helper) that calls
this.embed(text, embeddingModelName) with a sensible concurrency limit to avoid
request spikes and timeouts; keep logging and return shape the same but ensure
you await the concurrency-limited mapping so embeddings are returned as
number[][].

} catch (error) {
logger.error("Batch embedding generation failed", {
error: error instanceof Error ? error.message : String(error),
model: embeddingModelName,
count: texts.length,
});

throw this.handleProviderError(error);
}
}
}
Loading
Loading