diff --git a/README.md b/README.md index 3ff6ec9c1..948e06f34 100644 --- a/README.md +++ b/README.md @@ -47,7 +47,7 @@ const decision = await pipe.tryDecide({ **NeuroLink is the pipe layer of an AI nervous system.** Providers — OpenAI, Anthropic, Google, AWS, Azure, Mistral, local runtimes like Ollama, and dozens more — are the neurons: each generates a different kind of intelligence, at a different cost and latency. NeuroLink is the vascular layer that carries that intelligence, as a stream, to the applications — the organs — that consume it, across three inference types: `generate` and `stream` produce text, `decide` produces a calibrated `boolean`/`choice`/`score` judgment instead. A curated model registry (64 models, 132 aliases) backs metadata, routing, and context-window checks out of the box, and hundreds more models are reachable through aggregator providers — 100+ via LiteLLM, 300+ via OpenRouter. -Extracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — OpenAI, Anthropic, Google, AWS Bedrock, Azure, a local runtime, or any provider you add. `decide` is the third inference type — a typed, calibrated judgment instead of text — for the model-routing and gating decisions `generate`/`stream` were never meant to make, powered by a purpose-built decision model (TypeSafe Jev) rather than a general-purpose LLM: routing decisions land in ~400ms for about $0.00002, instead of a full generation call. +Extracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — OpenAI, Anthropic, Google, AWS Bedrock, Azure, a local runtime, or any provider you add. `decide` is the third inference type — a typed, calibrated judgment instead of text — for the model-routing and gating decisions `generate`/`stream` were never meant to make, powered by a purpose-built decision model (TypeSafe Jev, or the open-weights Laya) rather than a general-purpose LLM: with Jev, routing decisions land in ~400ms for about $0.00002, instead of a full generation call. **Why NeuroLink?** Three genuine inference types, not one dressed up three ways — `generate` and `stream` produce text; `decide` produces a calibrated `boolean`/`choice`/`score` judgment, and which types a provider serves is declared per-provider via `inferenceKinds` rather than inferred from behavior. Every neuron plugs into the same pipe, including 3 fully local runtimes (Ollama, LM Studio, llama.cpp) with per-request credential overrides, and MCP support covers all 4 transports (stdio, HTTP, SSE, WebSocket). Every AI-driven optimization the pipe performs — model routing, context compaction, tool selection — fails open: no key configured behaves exactly like NeuroLink without it, and routing uses asymmetric confidence thresholds (upgrade at 0.3, downgrade at 0.6) rather than a single cutoff, because a wrong downgrade costs more than a wrong upgrade. Switch providers with a single parameter change, leverage built-in tools plus any MCP-compliant tool server, deploy with confidence using enterprise features like Redis memory and multi-provider failover, and optimize costs automatically with intelligent routing. Use it via our professional CLI or TypeScript SDK—whichever fits your workflow. @@ -59,37 +59,37 @@ Extracted from production systems at Juspay, NeuroLink provides a practical, Typ ## What's New -| Feature | Version | Description | Guide | -| ------------------------------------------------------- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | -| **`decide` Inference Type + TypeSafe Jev** | next | A third inference type alongside `generate`/`stream`: typed, calibrated judgments (`boolean`, `choice`, `score`) via `neurolink.decide()` / `tryDecide()`, one parallel pass, ~400ms and ~$0.00002/decision. First provider is TypeSafe Jev (`TYPESAFE_API_KEY`, also reachable via the Vercel AI Gateway). Used internally for model routing, context budgeting, relevance compaction and tool routing — fail-open and a no-op without a key. Per-query RAG planning is opt-in via `RAGPipeline`. | [Decide Guide](docs/features/decide-inference-type.md) | -| **7 More Catalog Providers** | v12.11.0–v12.16.0 | Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage and API Route onboarded as Tier-2 catalog entries — one JSON file each, roster live-verified against the provider's own `/v1/models`. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) | -| **Claude-on-Vertex Proxy Fallback** | v12.18.0 | The Anthropic proxy pool can fall back to Claude served on Google Vertex, so an agentic turn survives losing its primary backend mid-conversation instead of failing the turn. | [Claude Proxy](docs/features/claude-proxy.md) | -| **Native-Loop V3 Conversation Reclaim** | v12.17.0 | Reclaims V3 conversations without splitting tool-call/tool-result pairs — the pairing a provider rejects the whole request over. | [Claude Proxy Architecture](docs/features/claude-proxy-architecture.md) | -| **Multi-Modal Embeddings** | v12.15.0 | `embed()` / `embedMany()` accept images alongside text on providers whose embedding models are multi-modal, for cross-modal retrieval in RAG and custom vector search. | [Embeddings Guide](docs/features/embeddings.md) | -| **Grok Build Auto-Configuration** | v12.14.0 | The proxy configures Grok Build automatically, deriving context windows and backends from the model catalog rather than hardcoded values. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) | -| **Anthropic Execution-Control Contract** | v12.13.0 | Truthful stream termination plus an opt-in execution-control contract, so a stream that stopped early reports why instead of looking like a clean finish. | [Claude Proxy](docs/features/claude-proxy.md) | -| **Catalog Tool Declarations Honoured at Runtime** | v12.12.0 | A Tier-2 catalog entry declaring `tools: false` (e.g. Mancer) no longer has tools offered to it at runtime — the JSON declaration is enforced, not just documented. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) | -| **Artifact Stores: Redis, Custom, Range Reads, Search** | v12.10.0 | Artifacts can be backed by Redis or a custom store, read by byte range, and searched — instead of being held only in process memory. | [Claude Proxy](docs/features/claude-proxy.md) | -| **Local CLI Spend Reading** | v12.6.0–v12.9.0 | Reads token usage directly from other coding CLIs' own local stores — Cursor, Grok Build, Hermes Agent and three more — and names them in proxy traffic, so spend is attributed per client. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) | -| **Native OpenAI Audio Streaming** | v12.7.0 | OpenAI TTS audio streams natively rather than being buffered to completion first. | [TTS Guide](docs/features/tts.md) | -| **HITL Pending-Confirmation State** | v12.5.0 | Exposes whether a human-in-the-loop confirmation is still outstanding, so a caller can distinguish 'waiting on a human' from 'finished'. | [Task Manager](docs/features/task-manager.md) | -| **OpenCode + Gemini CLI Proxy Clients** | v12.4.0 | OpenCode's generated config is actually loadable, and Gemini CLI is onboarded as a proxy client. | [OpenCode Proxy](docs/features/opencode-proxy-support.md) \| [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) | -| **SambaNova Provider** | v12.3.0 | RDU-accelerated open-weight flagships: Llama 3.3 70B (default), GPT-OSS 120B, DeepSeek V3.x, MiniMax, Gemma 4 (vision) — OpenAI-compatible Tier 2 catalog entry. Note: new SambaNova accounts require purchased credits. | [SambaNova Guide](docs/getting-started/providers/sambanova.md) | -| **Cerebras Provider** | v12.1.0 | Wafer-scale inference at ~3000 tok/s: GPT-OSS 120B (default) + Gemma 4 31B, OpenAI-compatible Tier 2 catalog entry, live-verified end to end (generate, stream, tools, structured output). | [Cerebras Guide](docs/getting-started/providers/cerebras.md) | -| **Avatar / Music Modalities + 12 Providers** | v9.65.0 | New `output: { mode: "avatar" \| "music" }` dispatch with handlers for D-ID, HeyGen, Replicate-MuseTalk (avatar) and Beatoven, ElevenLabs Music, Lyria, Replicate-MusicGen (music). Plus Fish Audio TTS, Kling/Runway/Replicate video, xAI/Groq/Cohere/Together/Fireworks/Perplexity/Cloudflare LLMs, Voyage/Jina embeddings, Stability/Ideogram/Recraft/Replicate image-gen. | [Provider Integration](docs/provider-integration/) | -| **Multi-Provider Voice (TTS/STT)** | v9.62.0 | 6 TTS providers (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Fish Audio, Cartesia) + 4 STT providers (Whisper, Deepgram, Azure STT, Google STT) + 2 realtime APIs (OpenAI Realtime, Gemini Live). | [TTS Guide](docs/features/tts.md) \| [STT Guide](docs/features/audio-input.md) \| [Realtime Guide](docs/features/real-time-services.md) | -| **4 New Providers** | v9.60.0 | DeepSeek (V3/R1), NVIDIA NIM (400+ catalog), LM Studio (local), llama.cpp (GGUF local). | [Provider Setup](docs/getting-started/provider-setup.md) | -| **ModelAccessDeniedError** | v9.59.0 | Typed `ModelAccessDeniedError` + `sdk.checkCredentials()` API for proactive credential validation before first call. | [Error Reference](docs/reference/troubleshooting.md) | -| **Provider Fallback Policy** | v9.58.0 | `providerFallback` callback + `modelChain` config for centralized multi-provider fallback logic. | [Advanced Guide](docs/advanced/index.md) | -| **Per-Request Credentials** | v9.52.0 | Pass credentials per-call or per-instance for all providers. Per-call overrides instance; instance overrides env vars. | [Credentials Guide](docs/features/per-request-credentials.md) | -| **AutoResearch** | v9.53.0 | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics — unattended for hours. | [AutoResearch Guide](docs/features/autoresearch.md) | -| **Gemini 3 Multi-turn Tool Fix** | v9.49.0 | Fixed multi-step agentic tool calling on Vertex AI Gemini 3. Correct `thoughtSignature` replay, `stepIndex` grouping, `executionId` session isolation, 5-min timeout. | [Vertex AI Guide](docs/getting-started/providers/google-vertex.md) | -| **MCP Enhancements** | v9.16.0 | Tool routing (6 strategies), result caching (LRU/FIFO/LFU), request batching, annotations, elicitation protocol, multi-server management. | [MCP Enhancements Guide](docs/features/mcp-enhancements.md) | -| **Memory** | v9.12.0 | Per-user condensed memory across conversations. LLM-powered condensation with S3, Redis, or SQLite. | [Memory Guide](docs/features/memory.md) | -| **Context Window Management** | v9.2.0 | 5-stage compaction pipeline with budget gate at 80% usage, per-provider token estimation. | [Context Compaction Guide](docs/features/context-compaction.md) | -| **Tool Execution Control** | v9.3.0 | `prepareStep` and `toolChoice` for per-step tool enforcement in multi-step agentic loops. | [API Reference](docs/api/type-aliases/GenerateOptions.md#preparestep) | -| **File Processor System** | v9.1.0 | 17+ file type processors with ProcessorRegistry, security sanitization, SVG text injection. | [File Processors Guide](docs/features/file-processors.md) | -| **RAG with generate()/stream()** | v9.2.0 | Pass `rag: { files }` for automatic document chunking, embedding, and AI-powered search. 10 chunking strategies, hybrid search, reranking, and a choice of 4 vector stores (in-memory, Chroma, PgVector, Pinecone). | [RAG Guide](docs/features/rag.md) | +| Feature | Version | Description | Guide | +| ------------------------------------------------------- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | +| **`decide` Inference Type + TypeSafe Jev + Laya** | next | A third inference type alongside `generate`/`stream`: typed, calibrated judgments (`boolean`, `choice`, `score`) via `neurolink.decide()` / `tryDecide()`, one parallel pass (~400ms and ~$0.00002/decision on Jev). First provider is TypeSafe Jev (`TYPESAFE_API_KEY`, also reachable via the Vercel AI Gateway); [Laya](docs/getting-started/providers/laya.md) (`LAYA_API_KEY` + `LAYA_BASE_URL`), an open-weights model you run yourself, is the second — TypeSafe wins when both are configured. Used internally for model routing, context budgeting, relevance compaction and tool routing — fail-open and a no-op without a key. Per-query RAG planning is opt-in via `RAGPipeline`. | [Decide Guide](docs/features/decide-inference-type.md) | +| **7 More Catalog Providers** | v12.11.0–v12.16.0 | Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage and API Route onboarded as Tier-2 catalog entries — one JSON file each, roster live-verified against the provider's own `/v1/models`. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) | +| **Claude-on-Vertex Proxy Fallback** | v12.18.0 | The Anthropic proxy pool can fall back to Claude served on Google Vertex, so an agentic turn survives losing its primary backend mid-conversation instead of failing the turn. | [Claude Proxy](docs/features/claude-proxy.md) | +| **Native-Loop V3 Conversation Reclaim** | v12.17.0 | Reclaims V3 conversations without splitting tool-call/tool-result pairs — the pairing a provider rejects the whole request over. | [Claude Proxy Architecture](docs/features/claude-proxy-architecture.md) | +| **Multi-Modal Embeddings** | v12.15.0 | `embed()` / `embedMany()` accept images alongside text on providers whose embedding models are multi-modal, for cross-modal retrieval in RAG and custom vector search. | [Embeddings Guide](docs/features/embeddings.md) | +| **Grok Build Auto-Configuration** | v12.14.0 | The proxy configures Grok Build automatically, deriving context windows and backends from the model catalog rather than hardcoded values. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) | +| **Anthropic Execution-Control Contract** | v12.13.0 | Truthful stream termination plus an opt-in execution-control contract, so a stream that stopped early reports why instead of looking like a clean finish. | [Claude Proxy](docs/features/claude-proxy.md) | +| **Catalog Tool Declarations Honoured at Runtime** | v12.12.0 | A Tier-2 catalog entry declaring `tools: false` (e.g. Mancer) no longer has tools offered to it at runtime — the JSON declaration is enforced, not just documented. | [Tier 2 Onboarding](docs/provider-integration/tiers/tier-2-catalog-entry.md) | +| **Artifact Stores: Redis, Custom, Range Reads, Search** | v12.10.0 | Artifacts can be backed by Redis or a custom store, read by byte range, and searched — instead of being held only in process memory. | [Claude Proxy](docs/features/claude-proxy.md) | +| **Local CLI Spend Reading** | v12.6.0–v12.9.0 | Reads token usage directly from other coding CLIs' own local stores — Cursor, Grok Build, Hermes Agent and three more — and names them in proxy traffic, so spend is attributed per client. | [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) | +| **Native OpenAI Audio Streaming** | v12.7.0 | OpenAI TTS audio streams natively rather than being buffered to completion first. | [TTS Guide](docs/features/tts.md) | +| **HITL Pending-Confirmation State** | v12.5.0 | Exposes whether a human-in-the-loop confirmation is still outstanding, so a caller can distinguish 'waiting on a human' from 'finished'. | [Task Manager](docs/features/task-manager.md) | +| **OpenCode + Gemini CLI Proxy Clients** | v12.4.0 | OpenCode's generated config is actually loadable, and Gemini CLI is onboarded as a proxy client. | [OpenCode Proxy](docs/features/opencode-proxy-support.md) \| [Proxy CLI Onboarding](docs/features/proxy-cli-onboarding.md) | +| **SambaNova Provider** | v12.3.0 | RDU-accelerated open-weight flagships: Llama 3.3 70B (default), GPT-OSS 120B, DeepSeek V3.x, MiniMax, Gemma 4 (vision) — OpenAI-compatible Tier 2 catalog entry. Note: new SambaNova accounts require purchased credits. | [SambaNova Guide](docs/getting-started/providers/sambanova.md) | +| **Cerebras Provider** | v12.1.0 | Wafer-scale inference at ~3000 tok/s: GPT-OSS 120B (default) + Gemma 4 31B, OpenAI-compatible Tier 2 catalog entry, live-verified end to end (generate, stream, tools, structured output). | [Cerebras Guide](docs/getting-started/providers/cerebras.md) | +| **Avatar / Music Modalities + 12 Providers** | v9.65.0 | New `output: { mode: "avatar" \| "music" }` dispatch with handlers for D-ID, HeyGen, Replicate-MuseTalk (avatar) and Beatoven, ElevenLabs Music, Lyria, Replicate-MusicGen (music). Plus Fish Audio TTS, Kling/Runway/Replicate video, xAI/Groq/Cohere/Together/Fireworks/Perplexity/Cloudflare LLMs, Voyage/Jina embeddings, Stability/Ideogram/Recraft/Replicate image-gen. | [Provider Integration](docs/provider-integration/) | +| **Multi-Provider Voice (TTS/STT)** | v9.62.0 | 6 TTS providers (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Fish Audio, Cartesia) + 4 STT providers (Whisper, Deepgram, Azure STT, Google STT) + 2 realtime APIs (OpenAI Realtime, Gemini Live). | [TTS Guide](docs/features/tts.md) \| [STT Guide](docs/features/audio-input.md) \| [Realtime Guide](docs/features/real-time-services.md) | +| **4 New Providers** | v9.60.0 | DeepSeek (V3/R1), NVIDIA NIM (400+ catalog), LM Studio (local), llama.cpp (GGUF local). | [Provider Setup](docs/getting-started/provider-setup.md) | +| **ModelAccessDeniedError** | v9.59.0 | Typed `ModelAccessDeniedError` + `sdk.checkCredentials()` API for proactive credential validation before first call. | [Error Reference](docs/reference/troubleshooting.md) | +| **Provider Fallback Policy** | v9.58.0 | `providerFallback` callback + `modelChain` config for centralized multi-provider fallback logic. | [Advanced Guide](docs/advanced/index.md) | +| **Per-Request Credentials** | v9.52.0 | Pass credentials per-call or per-instance for all providers. Per-call overrides instance; instance overrides env vars. | [Credentials Guide](docs/features/per-request-credentials.md) | +| **AutoResearch** | v9.53.0 | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics — unattended for hours. | [AutoResearch Guide](docs/features/autoresearch.md) | +| **Gemini 3 Multi-turn Tool Fix** | v9.49.0 | Fixed multi-step agentic tool calling on Vertex AI Gemini 3. Correct `thoughtSignature` replay, `stepIndex` grouping, `executionId` session isolation, 5-min timeout. | [Vertex AI Guide](docs/getting-started/providers/google-vertex.md) | +| **MCP Enhancements** | v9.16.0 | Tool routing (6 strategies), result caching (LRU/FIFO/LFU), request batching, annotations, elicitation protocol, multi-server management. | [MCP Enhancements Guide](docs/features/mcp-enhancements.md) | +| **Memory** | v9.12.0 | Per-user condensed memory across conversations. LLM-powered condensation with S3, Redis, or SQLite. | [Memory Guide](docs/features/memory.md) | +| **Context Window Management** | v9.2.0 | 5-stage compaction pipeline with budget gate at 80% usage, per-provider token estimation. | [Context Compaction Guide](docs/features/context-compaction.md) | +| **Tool Execution Control** | v9.3.0 | `prepareStep` and `toolChoice` for per-step tool enforcement in multi-step agentic loops. | [API Reference](docs/api/type-aliases/GenerateOptions.md#preparestep) | +| **File Processor System** | v9.1.0 | 17+ file type processors with ProcessorRegistry, security sanitization, SVG text injection. | [File Processors Guide](docs/features/file-processors.md) | +| **RAG with generate()/stream()** | v9.2.0 | Pass `rag: { files }` for automatic document chunking, embedding, and AI-powered search. 10 chunking strategies, hybrid search, reranking, and a choice of 4 vector stores (in-memory, Chroma, PgVector, Pinecone). | [RAG Guide](docs/features/rag.md) | ```typescript // decide() — a third inference type: calibrated judgments, not text (next) @@ -221,8 +221,8 @@ const neurolink = new NeuroLink({ ## Decide: Calibrated Judgments, Not Text -**NeuroLink now supports decision models — a third inference type, and the first -provider for it ships today.** +**NeuroLink now supports decision models — a third inference type, with two +providers today: a hosted one and an open-weights one you can run yourself.** `decide` sits alongside `generate` and `stream`. Instead of tokens, a decision model takes one `state` plus a map of named typed questions and returns one @@ -270,6 +270,20 @@ through the **Vercel AI Gateway** via `AI_GATEWAY_API_KEY`; force one transport with `TYPESAFE_TRANSPORT=direct|gateway`. Per-request credentials work as they do for every other provider. +[**Laya**](docs/getting-started/providers/laya.md) is the second decision +provider — Convai Innovations' Apache-2.0, open-weights "System One" model, for +when you'd rather run the decision model on your own infrastructure than call a +hosted one. There is no built-in endpoint: set `LAYA_API_KEY` and +`LAYA_BASE_URL` to point at a Laya server you run yourself, or a LiteLLM proxy +with a pass-through route to one (or pass `credentials.laya` to +`new NeuroLink({ credentials })`, or per call). Its encoders read a much +shorter state than TypeSafe's — about 768 tokens on the default +`typed-decisions` checkpoint (320 on `english`/`auto`), against TypeSafe's +~33,000 — so it fits short, structured decisions rather than long context. +Every built-in consumer below asks for the first configured decision provider, +and TypeSafe is listed first: with both configured, TypeSafe runs; with only +Laya's key and base URL set, Laya runs. + ```typescript import { NeuroLink, @@ -350,7 +364,7 @@ Measured against the live API — don't extrapolate past these: - Two input ceilings: `state` + the longest single question ≈ 33,000 tokens; `state` + all questions ≈ 64,000 tokens. - **Accuracy is the trade-off**: ~68% on TypeSafe's own 711-case benchmark vs. ~73% for a frontier model (TypeSafe's published figures, not our measurement). Use it for decisions that are **gated and reversible** — routing, dropping, budgeting — never for a final answer a user will see. -**[Decide Guide](docs/features/decide-inference-type.md)** · **[TypeSafe Provider Guide](docs/getting-started/providers/typesafe.md)** +**[Decide Guide](docs/features/decide-inference-type.md)** · **[TypeSafe Provider Guide](docs/getting-started/providers/typesafe.md)** · **[Laya Provider Guide](docs/getting-started/providers/laya.md)** ## Enterprise Security: Human-in-the-Loop (HITL) @@ -629,7 +643,7 @@ NeuroLink is a comprehensive AI development platform. Every feature below is shi ### 🤖 AI Provider Integration -**Every provider neuron behind one API** - Switch providers with a single parameter change. Nearly all serve `generate`/`stream`; TypeSafe Jev alone serves `decide`. Tool support: 29 native tool-calling, 3 model-dependent, 8 that serve no tools at all (embedding-, media- and decision-only). 3 are fully local runtimes (Ollama, LM Studio, llama.cpp) and 4 need zero configuration to start (those three plus LiteLLM) — no cloud account, no API key. 9 providers (OpenAI, Google AI Studio, Google Vertex, Amazon Bedrock, Cohere, Ollama, LiteLLM, Voyage, Jina) expose `embed()`/`embedMany()` natively for RAG and custom vector search. +**Every provider neuron behind one API** - Switch providers with a single parameter change. Nearly all serve `generate`/`stream`; TypeSafe Jev and Laya serve `decide` instead. Tool support: 30 native tool-calling, 3 model-dependent, 9 that serve no tools at all (embedding-, media- and decision-only). 3 are fully local runtimes (Ollama, LM Studio, llama.cpp) and 4 need zero configuration to start (those three plus LiteLLM) — no cloud account, no API key. 9 providers (OpenAI, Google AI Studio, Google Vertex, Amazon Bedrock, Cohere, Ollama, LiteLLM, Voyage, Jina) expose `embed()`/`embedMany()` natively for RAG and custom vector search. | Provider | Models | Free Tier | Tool Support | Status | Documentation | | --------------------- | -------------------------------------------------------------------------- | --------------- | ------------ | ------------- | ----------------------------------------------------------------------------------------------------------------------------- | @@ -663,9 +677,9 @@ NeuroLink is a comprehensive AI development platform. Every feature below is shi **Media generation** — [Replicate](docs/getting-started/providers/replicate.md) (`REPLICATE_API_TOKEN`) · [Stability AI](docs/getting-started/providers/stability.md) (`STABILITY_API_KEY`) · [Ideogram](docs/getting-started/providers/ideogram.md) (`IDEOGRAM_API_KEY`) · [Recraft](docs/getting-started/providers/recraft.md) (`RECRAFT_API_KEY`) -**Decision** — [TypeSafe Jev](docs/getting-started/providers/typesafe.md) (`TYPESAFE_API_KEY`, or `AI_GATEWAY_API_KEY` via the Vercel AI Gateway) — the only provider serving `decide` rather than `generate`/`stream`. +**Decision** — [TypeSafe Jev](docs/getting-started/providers/typesafe.md) (`TYPESAFE_API_KEY`, or `AI_GATEWAY_API_KEY` via the Vercel AI Gateway) · [Laya](docs/getting-started/providers/laya.md) (`LAYA_API_KEY` + `LAYA_BASE_URL`, open-weights, self-hosted) — the two providers serving `decide` rather than `generate`/`stream`. -**Decision-only provider:** **TypeSafe Jev** (`TYPESAFE_API_KEY`) does not appear in the table above because it does not serve `generate`/`stream` — it is the first provider for the `decide` inference type. See [Decide: Calibrated Judgments, Not Text](#decide-calibrated-judgments-not-text). +**Decision-only providers:** **TypeSafe Jev** (`TYPESAFE_API_KEY`) and **Laya** (`LAYA_API_KEY` + `LAYA_BASE_URL`) do not appear in the table above because neither serves `generate`/`stream` — they are the two providers for the `decide` inference type, TypeSafe first when both are configured. See [Decide: Calibrated Judgments, Not Text](#decide-calibrated-judgments-not-text). **[📖 Provider Comparison Guide](docs/reference/provider-comparison.md)** - Detailed feature matrix and selection criteria **[🔬 Provider Feature Compatibility](docs/reference/provider-feature-compatibility.md)** - Test-based compatibility reference for all 19 features across every provider @@ -1187,18 +1201,18 @@ Full command and API breakdown lives in [`docs/cli/commands.md`](docs/cli/comman ## Platform Capabilities at a Glance -| Capability | Highlights | -| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| **Provider unification** | Every provider neuron behind one API, with automatic fallback, cost-aware routing, `providerFallback` policy, `modelChain` config. | -| **Decision inference** | Third inference type (`decide`) alongside generate/stream: calibrated `boolean`/`choice`/`score` judgments via TypeSafe Jev, ~400ms flat, ~$0.00002/decision. Used internally for model routing, context budgeting, relevance compaction and tool routing; per-query RAG planning is opt-in via `RAGPipeline`. | -| **Multimodal pipeline** | Stream images + CSV data + PDF documents across providers with local/remote assets. Auto-detection for mixed file types. | -| **Voice pipeline** | TTS (6 providers: Google, OpenAI, ElevenLabs, Azure, Fish Audio, Cartesia) + STT (4 providers) + realtime voice APIs (OpenAI Realtime, Gemini Live). | -| **Quality & governance** | Auto-evaluation engine (14 scorers), guardrails middleware, HITL workflows, audit logging. | -| **Memory & context** | Per-user condensed memory (S3/Redis/SQLite), Redis session export, 5-stage context compaction. | -| **CLI tooling** | 34 commands: loop sessions, setup wizard, config validation, Redis auto-detect, JSON output, TTS/STT flags. | -| **Enterprise ops** | Claude proxy, OTLP observability, OpenObserve dashboard, regional routing, credential management. | -| **Tool ecosystem** | MCP auto discovery, HTTP/stdio/SSE/WebSocket transports, LiteLLM hub access, SageMaker custom deployment, web search. | -| **Engineering rigor** | 129 end-to-end test suites (every suite drives the public `generate`/`stream`/`decide`/CLI surface, never internals), 13 custom ESLint rules enforcing the architecture (no `interface`, unique type names, barrel-only type imports) — all AST-based, no regex heuristics. | +| Capability | Highlights | +| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **Provider unification** | Every provider neuron behind one API, with automatic fallback, cost-aware routing, `providerFallback` policy, `modelChain` config. | +| **Decision inference** | Third inference type (`decide`) alongside generate/stream: calibrated `boolean`/`choice`/`score` judgments via TypeSafe Jev (~400ms flat, ~$0.00002/decision) or Laya, a self-hosted open-weights alternative. Used internally for model routing, context budgeting, relevance compaction and tool routing; per-query RAG planning is opt-in via `RAGPipeline`. | +| **Multimodal pipeline** | Stream images + CSV data + PDF documents across providers with local/remote assets. Auto-detection for mixed file types. | +| **Voice pipeline** | TTS (6 providers: Google, OpenAI, ElevenLabs, Azure, Fish Audio, Cartesia) + STT (4 providers) + realtime voice APIs (OpenAI Realtime, Gemini Live). | +| **Quality & governance** | Auto-evaluation engine (14 scorers), guardrails middleware, HITL workflows, audit logging. | +| **Memory & context** | Per-user condensed memory (S3/Redis/SQLite), Redis session export, 5-stage context compaction. | +| **CLI tooling** | 34 commands: loop sessions, setup wizard, config validation, Redis auto-detect, JSON output, TTS/STT flags. | +| **Enterprise ops** | Claude proxy, OTLP observability, OpenObserve dashboard, regional routing, credential management. | +| **Tool ecosystem** | MCP auto discovery, HTTP/stdio/SSE/WebSocket transports, LiteLLM hub access, SageMaker custom deployment, web search. | +| **Engineering rigor** | 129 end-to-end test suites (every suite drives the public `generate`/`stream`/`decide`/CLI surface, never internals), 13 custom ESLint rules enforcing the architecture (no `interface`, unique type names, barrel-only type imports) — all AST-based, no regex heuristics. | ## Documentation Map @@ -1226,7 +1240,7 @@ Full command and API breakdown lives in [`docs/cli/commands.md`](docs/cli/comman **Decision Inference:** -- [Decide Guide](docs/features/decide-inference-type.md) - The `decide` inference type: boolean/choice/score primitives, TypeSafe Jev setup, measured latency/cost +- [Decide Guide](docs/features/decide-inference-type.md) - The `decide` inference type: boolean/choice/score primitives, TypeSafe Jev / Laya setup, measured latency/cost **Provider Intelligence:** diff --git a/docs-site/src/pages/index.tsx b/docs-site/src/pages/index.tsx index 1a1cc23da..3cda0355e 100644 --- a/docs-site/src/pages/index.tsx +++ b/docs-site/src/pages/index.tsx @@ -27,6 +27,7 @@ const PROVIDERS = [ "Azure Speech", "OpenAI TTS", "TypeSafe", + "Laya", ]; const QUICK_LINKS = [ @@ -118,7 +119,7 @@ const FAQ_ITEMS = [ { question: "What is NeuroLink?", answer: - "NeuroLink is an enterprise AI development platform that provides unified access to 40 AI providers (OpenAI, Anthropic, Google AI, AWS Bedrock, Azure, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus voice providers like ElevenLabs, Deepgram, and more) through a single TypeScript SDK and professional CLI — 39 of them for generate()/stream(), plus TypeSafe for calibrated decide() judgements. It is extracted from production systems at Juspay and battle-tested at enterprise scale.", + "NeuroLink is an enterprise AI development platform that provides unified access to AI providers (OpenAI, Anthropic, Google AI, AWS Bedrock, Azure, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus voice providers like ElevenLabs, Deepgram, and more) through a single TypeScript SDK and professional CLI — most of them for generate()/stream(), plus TypeSafe and Laya for calibrated decide() judgements. It is extracted from production systems at Juspay and battle-tested at enterprise scale.", }, { question: "How is NeuroLink different from LangChain or Vercel AI SDK?", @@ -133,7 +134,7 @@ const FAQ_ITEMS = [ { question: "What AI providers does NeuroLink support?", answer: - "NeuroLink supports 40 providers in total. 39 serve generate()/stream(): OpenAI, Anthropic, Google AI Studio, Google Vertex AI, AWS Bedrock, Azure OpenAI, Mistral, Ollama, LiteLLM, HuggingFace, SageMaker, OpenRouter, DeepSeek, NVIDIA NIM (400+ catalog models), LM Studio (local), llama.cpp (local GGUF), any OpenAI-compatible endpoint, and voice providers — OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT. The 40th, TypeSafe, serves only decide() — calibrated typed judgements, no text. Switching generate/stream providers requires changing a single parameter.", + "Most NeuroLink providers serve generate()/stream(): OpenAI, Anthropic, Google AI Studio, Google Vertex AI, AWS Bedrock, Azure OpenAI, Mistral, Ollama, LiteLLM, HuggingFace, SageMaker, OpenRouter, DeepSeek, NVIDIA NIM (400+ catalog models), LM Studio (local), llama.cpp (local GGUF), any OpenAI-compatible endpoint, and voice providers — OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT. Two serve decide() only — calibrated typed judgements, no text: TypeSafe, and Laya, a self-hosted open-weights model (LAYA_API_KEY + LAYA_BASE_URL). Switching generate/stream providers requires changing a single parameter.", }, { question: "Does NeuroLink support MCP (Model Context Protocol)?", @@ -148,7 +149,7 @@ const FAQ_ITEMS = [ { question: "What is the `decide` inference type?", answer: - "decide() is a third inference type alongside generate() and stream() — a separate modality, not a mode of the other two, so a text-only provider is never reachable from it and a decide-only provider is never reachable from generation fallback. A decision model takes one state plus a set of named, typed questions and returns one typed, calibrated answer per question in a single parallel pass — no text, nothing to parse out of prose. Because the confidence is calibrated rather than self-reported, it can gate action directly: NeuroLink's classifier router uses asymmetric thresholds (0.3 confidence to route a request up to a pricier model, 0.6 to route it down) since routing too cheap and routing too expensive don't cost the same. Six consumers build on it today — model routing, a registry-derived model catalogue (64 models across 7 providers, 132 aliases, added on top of a host's own declared pool rather than replacing it), context budget sizing, compaction quality gating, calibrated MCP/tool-server routing, and RAG retrieval planning — and every one fails open: with no TYPESAFE_API_KEY configured, behavior is byte-for-byte what it was before. Decisions get their own model.decision observability span with independent cost attribution, kept separate from generation metrics, across all 9 supported exporters. TypeSafe's Jev is the first decision provider; enabling it is that one environment variable.", + "decide() is a third inference type alongside generate() and stream() — a separate modality, not a mode of the other two, so a text-only provider is never reachable from it and a decide-only provider is never reachable from generation fallback. A decision model takes one state plus a set of named, typed questions and returns one typed, calibrated answer per question in a single parallel pass — no text, nothing to parse out of prose. Because the confidence is calibrated rather than self-reported, it can gate action directly: NeuroLink's classifier router uses asymmetric thresholds (0.3 confidence to route a request up to a pricier model, 0.6 to route it down) since routing too cheap and routing too expensive don't cost the same. Six consumers build on it today — model routing, a registry-derived model catalogue (64 models across 7 providers, 132 aliases, added on top of a host's own declared pool rather than replacing it), context budget sizing, compaction quality gating, calibrated MCP/tool-server routing, and RAG retrieval planning — and every one fails open: with no TYPESAFE_API_KEY configured, behavior is byte-for-byte what it was before. Decisions get their own model.decision observability span with independent cost attribution, kept separate from generation metrics, across all 9 supported exporters. TypeSafe's Jev is the first decision provider, and Laya — an open-weights model you can run yourself — is a second; enabling either is one environment variable (two for Laya, which also needs a base URL).", }, ]; diff --git a/docs-site/static/search-index.json b/docs-site/static/search-index.json index 7ab137fc0..d55fd9623 100644 --- a/docs-site/static/search-index.json +++ b/docs-site/static/search-index.json @@ -113,9 +113,9 @@ {"objectID":"b38e417fcd8ea5d3b983875efebb5b5a309883fdfe8bfdc80300e473bea5e081","title":"Gradual Adoption","url":"/docs/WORKFLOW-ENGINE-LLD#gradual-adoption","content":"Phase 1: Users can try workflows alongside existing methods\nPhase 2: Workflows become recommended for high-stakes queries\nPhase 3: Workflows are default with single-model as fallback","hierarchy":{"lvl0":"WORKFLOW ENGINE LLD","lvl1":"Workflow Engine - Low-Level Design","lvl2":"Gradual Adoption","lvl3":""}}, {"objectID":"2127ec4796813559375c700e37b4944f77c72f22a0410fe93bd6f5f7cff8a4c0","title":"17. Performance Benchmarks (Expected)","url":"/docs/WORKFLOW-ENGINE-LLD#17-performance-benchmarks-expected","content":"| Workflow | Models | Judge | Latency (p50) | Latency (p95) | Cost Multiplier |\n| ------------- | ------ | ----- | ------------- | ------------- | --------------- |\n| consensus-3 | 3 | 1 | 3.2s | 5.1s | 4.2x |\n| fast-fallback | 1-2 | 0 | 1.1s | 2.8s | 1.3x |\n| quality-max | 2 | 1 | 3.5s | 4.9s | 3.1x |\n| multi-judge-5 | 3 | 2 | 4.8s | 6.7s | 5.3x |","hierarchy":{"lvl0":"WORKFLOW ENGINE LLD","lvl1":"Workflow Engine - Low-Level Design","lvl2":"17. Performance Benchmarks (Expected)","lvl3":""}}, {"objectID":"2f54303255393f1b5c8695c6342964ce83cb786fd3ae00f535bf17c44cae16cd","title":"📝 Implementation Checklist","url":"/docs/WORKFLOW-ENGINE-LLD#-implementation-checklist","content":"[ ] Create directory structure\n[ ] Implement with all interfaces\n[ ] Implement with Zod schemas\n[ ] Implement \n[ ] Implement \n[ ] Implement \n[ ] Implement \n[ ] Implement \n[ ] Create built-in workflows (consensus, fallback, quality-max)\n[ ] Add methods to class\n[ ] Export types from \n[ ] Write unit tests (80% coverage target)\n[ ] Write integration tests\n[ ] Add JSDoc documentation\n[ ] Create user guide with examples\n[ ] Add CLI support (optional Phase 2)\n\nDocument Status: ✅ Ready for Implementation \nNext Step: Code generation upon approval","hierarchy":{"lvl0":"WORKFLOW ENGINE LLD","lvl1":"Workflow Engine - Low-Level Design","lvl2":"📝 Implementation Checklist","lvl3":""}}, -{"objectID":"de72b620d608319199e3abf25b54298afc8fe5f294c73c538956a2feb63fe9b1","title":"The Nervous System Model","url":"/docs/about/nervous-system-model","content":"The Nervous System Model\n\nNeuroLink is built around a biological metaphor — not as decoration, but as a structural model that governs every architectural decision.\n\nThe Three Components\n\nNeurons — LLM Providers\n\nNeurons are where intelligence is generated. In NeuroLink, neurons are the 40 AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.\n\nThe Pipe — NeuroLink\n\nThe pipe is the vascular layer that carries streams between neurons and organs. This is NeuroLink itself.\n\nWhat the pipe does every time you call or :\nContext Building — RAG retrieval, memory lookup, file processing merge into the prompt\nBudget Check — BudgetChecker validates the assembled context fits the model's window\nProvider Dispatch — ProviderRegistry routes to the correct neuron\nStream Emission — Tokens flow as an async iterable\nTool Interception — When the model calls a tool, the stream pauses, MCP tool executes, result injects, stream continues\nObservability — Every stage emits OpenTelemetry spans\n\nOrgans — Connectors\n\nOrgans are the applications that consume the pipe. They connect to the vascular layer and open a gateway — a specific way for people or systems to interact with AI.\n\nEvery application built on NeuroLink is an organ. Production organs today:\nAutomatic — Shopify operations hub: address intelligence, RTO risk scoring\nTara — Slack engineering assistant: conversational AI with MCP tool access\nYama — Code review judge: automated PR analysis and governance\n\nWhy This Model Works\n\nThe metaphor enforces good architecture:\n\nSeparation of concerns: Neurons (generation) and organs (consumption) are completely decoupled. Changing AI provider doesn't touch the application. Changing the application doesn't touch the provider.\n\nSingle flow direction: Intelligence flows one way — neuron → pipe → organ. There's no confusion about where logic lives.\n\nObservable by default: A vascular system you can't monitor is dangerous. Every stage of the pipe emits telemetry by design.\n\nComposable: Multiple organs can share the same pipe. One NeuroLink instance serves many connectors.\n\nExtending the System\n\nThe nervous system model scales in three directions:\nAdd neurons — New AI provider? Register it in ProviderRegistry.\nExtend the pipe — New capability (chunking strategy, reranker, compaction stage)? Add it to the pipeline.\nBuild organs — New application? Import NeuroLink, connect to the pipe, open your gateway.\n\nSee Pipe Architecture → for the technical implementation.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"","lvl3":""}}, +{"objectID":"de72b620d608319199e3abf25b54298afc8fe5f294c73c538956a2feb63fe9b1","title":"The Nervous System Model","url":"/docs/about/nervous-system-model","content":"The Nervous System Model\n\nNeuroLink is built around a biological metaphor — not as decoration, but as a structural model that governs every architectural decision.\n\nThe Three Components\n\nNeurons — LLM Providers\n\nNeurons are where intelligence is generated. In NeuroLink, neurons are the AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev and Laya (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.\n\nThe Pipe — NeuroLink\n\nThe pipe is the vascular layer that carries streams between neurons and organs. This is NeuroLink itself.\n\nWhat the pipe does every time you call or :\nContext Building — RAG retrieval, memory lookup, file processing merge into the prompt\nBudget Check — BudgetChecker validates the assembled context fits the model's window\nProvider Dispatch — ProviderRegistry routes to the correct neuron\nStream Emission — Tokens flow as an async iterable\nTool Interception — When the model calls a tool, the stream pauses, MCP tool executes, result injects, stream continues\nObservability — Every stage emits OpenTelemetry spans\n\nOrgans — Connectors\n\nOrgans are the applications that consume the pipe. They connect to the vascular layer and open a gateway — a specific way for people or systems to interact with AI.\n\nEvery application built on NeuroLink is an organ. Production organs today:\nAutomatic — Shopify operations hub: address intelligence, RTO risk scoring\nTara — Slack engineering assistant: conversational AI with MCP tool access\nYama — Code review judge: automated PR analysis and governance\n\nWhy This Model Works\n\nThe metaphor enforces good architecture:\n\nSeparation of concerns: Neurons (generation) and organs (consumption) are completely decoupled. Changing AI provider doesn't touch the application. Changing the application doesn't touch the provider.\n\nSingle flow direction: Intelligence flows one way — neuron → pipe → organ. There's no confusion about where logic lives.\n\nObservable by default: A vascular system you can't monitor is dangerous. Every stage of the pipe emits telemetry by design.\n\nComposable: Multiple organs can share the same pipe. One NeuroLink instance serves many connectors.\n\nExtending the System\n\nThe nervous system model scales in three directions:\nAdd neurons — New AI provider? Register it in ProviderRegistry.\nExtend the pipe — New capability (chunking strategy, reranker, compaction stage)? Add it to the pipeline.\nBuild organs — New application? Import NeuroLink, connect to the pipe, open your gateway.\n\nSee Pipe Architecture → for the technical implementation.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"","lvl3":""}}, {"objectID":"27ef114b7e0cecc6ca45dbf3e9285980b691cd27d21e1912f2a27a7f309f68fe","title":"The Nervous System Model","url":"/docs/about/nervous-system-model#the-nervous-system-model","content":"NeuroLink is built around a biological metaphor — not as decoration, but as a structural model that governs every architectural decision.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"The Nervous System Model","lvl3":""}}, -{"objectID":"58e5ab0de0df2a0949ae12798cd53ac3133b2c253880bcbf0c573938376fd84f","title":"Neurons — LLM Providers","url":"/docs/about/nervous-system-model#neurons-llm-providers","content":"Neurons are where intelligence is generated. In NeuroLink, neurons are the 40 AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Neurons — LLM Providers","lvl3":""}}, +{"objectID":"58e5ab0de0df2a0949ae12798cd53ac3133b2c253880bcbf0c573938376fd84f","title":"Neurons — LLM Providers","url":"/docs/about/nervous-system-model#neurons-llm-providers","content":"Neurons are where intelligence is generated. In NeuroLink, neurons are the AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev and Laya (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar).\n\nEach provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Neurons — LLM Providers","lvl3":""}}, {"objectID":"96f94c64476fe5567b61e5374da6d1b5c5020e22654c6c082b8936ae3f72eb51","title":"The Pipe — NeuroLink","url":"/docs/about/nervous-system-model#the-pipe-neurolink","content":"The pipe is the vascular layer that carries streams between neurons and organs. This is NeuroLink itself.\n\nWhat the pipe does every time you call or :\nContext Building — RAG retrieval, memory lookup, file processing merge into the prompt\nBudget Check — BudgetChecker validates the assembled context fits the model's window\nProvider Dispatch — ProviderRegistry routes to the correct neuron\nStream Emission — Tokens flow as an async iterable\nTool Interception — When the model calls a tool, the stream pauses, MCP tool executes, result injects, stream continues\nObservability — Every stage emits OpenTelemetry spans","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"The Pipe — NeuroLink","lvl3":""}}, {"objectID":"1c74200e8d4572e84dca981b50bb162912b4c6700ea7495d60fe286cfb16347f","title":"Organs — Connectors","url":"/docs/about/nervous-system-model#organs-connectors","content":"Organs are the applications that consume the pipe. They connect to the vascular layer and open a gateway — a specific way for people or systems to interact with AI.\n\nEvery application built on NeuroLink is an organ. Production organs today:\nAutomatic — Shopify operations hub: address intelligence, RTO risk scoring\nTara — Slack engineering assistant: conversational AI with MCP tool access\nYama — Code review judge: automated PR analysis and governance","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Organs — Connectors","lvl3":""}}, {"objectID":"6c97c4e15e0de12e62fbb1994cc7b18328c9f2b0afa2b721d8f40f01e70d32f7","title":"Why This Model Works","url":"/docs/about/nervous-system-model#why-this-model-works","content":"The metaphor enforces good architecture:\n\nSeparation of concerns: Neurons (generation) and organs (consumption) are completely decoupled. Changing AI provider doesn't touch the application. Changing the application doesn't touch the provider.\n\nSingle flow direction: Intelligence flows one way — neuron → pipe → organ. There's no confusion about where logic lives.\n\nObservable by default: A vascular system you can't monitor is dangerous. Every stage of the pipe emits telemetry by design.\n\nComposable: Multiple organs can share the same pipe. One NeuroLink instance serves many connectors.","hierarchy":{"lvl0":"About","lvl1":"The Nervous System Model","lvl2":"Why This Model Works","lvl3":""}}, @@ -3541,9 +3541,9 @@ {"objectID":"a6f80762fe08ca735f864be4541b6863065a0acfc5849344db2e40f38db28db3","title":"Verify file","url":"/docs/features/image-generation#verify-file","content":"file ./test-output/square.png","hierarchy":{"lvl0":"Features","lvl1":"Image Generation Streaming Guide","lvl2":"Verify file","lvl3":""}}, {"objectID":"0de2f95f5125f67748b214f7ed674edcf44078cc415fa081f6edf46428d7c65f","title":"Output: ./test-output/square.png: PNG image data, 1024 x 1024, 8-bit/color RGB","url":"/docs/features/image-generation#output-test-outputsquarepng-png-image-data-1024-x-1024-8-bitcolor-rgb","content":"`","hierarchy":{"lvl0":"Features","lvl1":"Image Generation Streaming Guide","lvl2":"Output: ./test-output/square.png: PNG image data, 1024 x 1024, 8-bit/color RGB","lvl3":""}}, {"objectID":"aa931e13d4db435aa9de30f72f599057c2977804e4316c3cb61f84d8e43fd97e","title":"Conclusion","url":"/docs/features/image-generation#conclusion","content":"NeuroLink's image generation streaming provides a unified interface for both text and image generation. The fake streaming approach ensures consistency while maintaining the benefits of streaming APIs. By following the patterns and examples in this guide, you can effectively integrate image generation into your applications.\n\nFor more information:\nAPI Reference\nProvider Comparison\nProvider Status Monitoring","hierarchy":{"lvl0":"Features","lvl1":"Image Generation Streaming Guide","lvl2":"Conclusion","lvl3":""}}, -{"objectID":"e3366024d1a12d0c73fbcb1592927d2dcc0d92c12eaea661cf5164bb7217c32f","title":"Feature Guides","url":"/docs/features","content":"Feature Guides\n\nComprehensive guides for all NeuroLink features organized by category. Each guide includes setup, usage patterns, configuration, and troubleshooting.\n\nLatest Features (Q1 2026)\n\n| Feature | Description |\n| ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| Inference Type | A third inference type alongside /: typed, calibrated // judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision. First provider is TypeSafe Jev. Fail-open: a no-op without a key. |\n| Model Routing with a Decision Model | One round trip answers difficulty, capabilities, risk, context scope and the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). |\n| Model Catalogue | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One question ranks all N candidates. |\n| Context Budget | A per-request compaction threshold derived from how much context the request actually needs. Only ever lowers the default, never raises it. |\n| Relevance Compaction | Stage 0 of the compaction pipeline: asks which earlier messages the current request still needs, plus a quality gate on the generated summary. |\n| Tool Routing with a Decision Model | One calibrated yes/no per MCP server, dropping only on a confident \"no\" — replaces a 15s LLM call at ~400ms. |\n| RAG Retrieval Planning | Per-query breadth and whether to use hybrid / graph / rerank. Opt-in via . |\n| Real-time Voice Services | Bidirectional realtime voice APIs — OpenAI Realtime and Gemini Live. Full-duplex audio streaming with tool calls, barge-in, and interruption. |\n| LiveKit Voice Agent | WebRTC voice agent using LiveKit for the real-time loop (transport, VAD, turn-taking, worker-per-call scaling) with NeuroLink as the brain (LLM, tools, memory). Cloud or self-hosted. |\n| Provider Fallback & Model Chains | callback + config (v9.58.0) — centralized multi-provider fallback policy for resilient AI workflows. |\n| Credential Validation | Pre-flight API + typed (v9.59.0) — actionable credential errors and validation before first call. |\n| AutoResearch | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics, keeps improvements — runs unattended for hours. |\n| MCP Enhancements | Advanced MCP features: ToolRouter, ToolCache, RequestBatcher, tool annotations, elicitation protocol, and custom MCP server creation. ","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"","lvl3":""}}, +{"objectID":"e3366024d1a12d0c73fbcb1592927d2dcc0d92c12eaea661cf5164bb7217c32f","title":"Feature Guides","url":"/docs/features","content":"Feature Guides\n\nComprehensive guides for all NeuroLink features organized by category. Each guide includes setup, usage patterns, configuration, and troubleshooting.\n\nLatest Features (Q1 2026)\n\n| Feature | Description |\n| ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Inference Type | A third inference type alongside /: typed, calibrated // judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision. First provider is TypeSafe Jev; Laya, open-weights and self-hosted, is the second. Fail-open: a no-op without a key. |\n| Model Routing with a Decision Model | One round trip answers difficulty, capabilities, risk, context scope and the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). |\n| Model Catalogue | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One question ranks all N candidates. |\n| Context Budget | A per-request compaction threshold derived from how much context the request actually needs. Only ever lowers the default, never raises it. |\n| Relevance Compaction | Stage 0 of the compaction pipeline: asks which earlier messages the current request still needs, plus a quality gate on the generated summary. |\n| Tool Routing with a Decision Model | One calibrated yes/no per MCP server, dropping only on a confident \"no\" — replaces a 15s LLM call at ~400ms. |\n| RAG Retrieval Planning | Per-query breadth and whether to use hybrid / graph / rerank. Opt-in via . |\n| Real-time Voice Services | Bidirectional realtime voice APIs — OpenAI Realtime and Gemini Live. Full-duplex audio streaming with tool calls, barge-in, and interruption. |\n| LiveKit Voice Agent | WebRTC voice agent using LiveKit for the real-time loop (transport, VAD, turn-taking, worker-per-call scaling) with NeuroLink as the brain (LLM, tools, memory). Cloud or self-hosted. |\n| Provider Fallback & Model Chains | callback + config (v9.58.0) — centralized multi-provider fallback policy for resilient AI workflows. ","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"","lvl3":""}}, {"objectID":"a477c76a577a17e7aeaf89736289a0a41734b4bcfe479a523cdd9e2051bd0ba7","title":"Feature Guides","url":"/docs/features#feature-guides","content":"Comprehensive guides for all NeuroLink features organized by category. Each guide includes setup, usage patterns, configuration, and troubleshooting.","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Feature Guides","lvl3":""}}, -{"objectID":"a43e81fa4a6c594f8b6b4cfc216db1356d75b3e4ebb54279dd2da7f69001d28f","title":"Latest Features (Q1 2026)","url":"/docs/features#latest-features-q1-2026","content":"| Feature | Description |\n| ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |\n| Inference Type | A third inference type alongside /: typed, calibrated // judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision. First provider is TypeSafe Jev. Fail-open: a no-op without a key. |\n| Model Routing with a Decision Model | One round trip answers difficulty, capabilities, risk, context scope and the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). |\n| Model Catalogue | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One question ranks all N candidates. |\n| Context Budget | A per-request compaction threshold derived from how much context the request actually needs. Only ever lowers the default, never raises it. |\n| Relevance Compacti","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Latest Features (Q1 2026)","lvl3":""}}, +{"objectID":"a43e81fa4a6c594f8b6b4cfc216db1356d75b3e4ebb54279dd2da7f69001d28f","title":"Latest Features (Q1 2026)","url":"/docs/features#latest-features-q1-2026","content":"| Feature | Description |\n| ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Inference Type | A third inference type alongside /: typed, calibrated // judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision. First provider is TypeSafe Jev; Laya, open-weights and self-hosted, is the second. Fail-open: a no-op without a key. |\n| Model Routing with a Decision Model | One round trip answers difficulty, capabilities, risk, context scope and the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). |\n| Model Catalogue | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One question ranks all N candidates. ","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Latest Features (Q1 2026)","lvl3":""}}, {"objectID":"84c604ce940f85fbe8aa8d94675e4c1dae80b22ede558d58102426598f6a7d06","title":"Core Features (shipped 2025)","url":"/docs/features#core-features-shipped-2025","content":"| Feature | Description |\n| -------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |\n| Image Generation | Generate images from text prompts using Gemini models via Vertex AI or Google AI Studio. |\n| Enterprise HITL | Production-ready HITL with approval workflows, confidence thresholds, and enterprise patterns. |\n| Interactive CLI | AI development environment with loop mode, session variables, and conversation memory. |\n| MCP Tools Showcase | Complete guide to 6 built-in tools and connecting external MCP servers across 6 categories. |\n| Human-in-the-Loop (HITL) | Pause AI tool execution for user approval before risky operations like file deletion or API calls. |\n| Guardrails Middleware | Content filtering, PII detection, and safety checks for AI outputs with zero configuration. |\n| Redis Conversation Export | Export complete session history as JSON for analytics, debugging, and compliance auditing. |\n| Context Compaction | Automatic conversation compression for long-running sessions to stay within token limits. |\n| LiteLLM Integration | Access 100+ AI models from all major providers through unified LiteLLM routing interface. |\n| SageMaker Integration | Deploy and use custom-trained models on AWS SageMaker infrastructure with full control. |","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Core Features (shipped 2025)","lvl3":""}}, {"objectID":"f64aba3e564a8f56064839c33b533971468da990a9cc95d86ceb1e9d13223478","title":"Earlier Core Features (shipped Q3 2025)","url":"/docs/features#earlier-core-features-shipped-q3-2025","content":"| Feature | Description |\n| ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |\n| Multimodal Chat Experiences | Stream text and images together with automatic provider fallbacks and format conversion. |\n| CSV File Support | Process CSV files for data analysis with automatic format conversion. Works with all providers. |\n| PDF File Support | Process PDF documents for visual analysis and content extraction. Native provider support. |\n| Office Documents | Process DOCX, PPTX, XLSX files for document analysis. Native Bedrock, Vertex, Anthropic support. |\n| Auto Evaluation Engine | Automated quality scoring and metrics export for AI response validation using LLM-as-judge. |\n| CLI Loop Sessions | Persistent interactive mode with conversation memory and session state for prompt engineering. |\n| Regional Streaming Controls | Region-specific model deployment and routing for compliance and latency optimization. |\n| Provider Orchestration Brain | Adaptive provider and model selection with intelligent fallbacks based on task classification. |","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Earlier Core Features (shipped Q3 2025)","lvl3":""}}, {"objectID":"fbd682fc148477639d38ab175c38363d003a1464834e84b2a48d88087d001b91","title":"Platform Capabilities at a Glance","url":"/docs/features#platform-capabilities-at-a-glance","content":"| Category | Features | Documentation |\n| ------------------------ | ------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- |\n| Provider unification | 40 providers with automatic failover, cost-aware routing, policy, config | Provider Setup |\n| Multimodal pipeline | Stream images + CSV data + PDF documents + Office files across providers with auto-detection for mixed file types. | Multimodal Guide, CSV Support, PDF Support, Office Docs |\n| Voice pipeline | TTS (4 providers) + STT (4 providers) + realtime APIs (OpenAI Realtime, Gemini Live) | TTS Guide, STT Guide, Realtime Services |\n| Quality & governance | Auto-evaluation engine (14 scorers), guardrails middleware, HITL workflows, audit logging | Auto Evaluation, Guardrails, HITL |\n| Memory & context | Per-user condensed memory (S3/Redis/SQLite), Redis session export, 5-stage context compaction | Conversation Memory, Memory, Redis Export |\n| CLI tooling | Loop sessions, setup wizard, config validation, Redis auto-detect, JSON output, TTS/STT flags | CLI Loop, CLI Commands |\n| Enterprise ops | Claude proxy, OTLP observability, OpenObserve dashboard, regional routing, credential management ","hierarchy":{"lvl0":"Features","lvl1":"Feature Guides","lvl2":"Platform Capabilities at a Glance","lvl3":""}}, @@ -4893,11 +4893,11 @@ {"objectID":"8bee50a94bc91cf5273efb43aa2f13a39be6cd72e048c55da55bd80bc74ce8a0","title":"Getting Help","url":"/docs/getting-started/installation#getting-help","content":"Check our Troubleshooting Guide\nReview FAQ\nSearch GitHub Issues\nCreate new issue with:\nNode.js version ()\nOperating system\nError message\nSteps to reproduce","hierarchy":{"lvl0":"Getting Started","lvl1":"Installation","lvl2":"Getting Help","lvl3":""}}, {"objectID":"80210a84a233c58e9a12aa96cedd225ad03c3bfe5344b78032d691be93ae3ba0","title":"Verification Checklist","url":"/docs/getting-started/installation#verification-checklist","content":"[ ] Node.js 18+ installed\n[ ] NeuroLink package installed or accessible via npx\n[ ] API keys configured in file\n[ ] shows working providers\n[ ] Basic generation command works\n[ ] TypeScript support (if needed)\n[ ] Framework integration (if applicable)","hierarchy":{"lvl0":"Getting Started","lvl1":"Installation","lvl2":"Verification Checklist","lvl3":""}}, {"objectID":"44c035a6ed0dc3b4a04ad68e17bc22de77840cb664eaee4e84f04bb8bd70caaf","title":"Next Steps","url":"/docs/getting-started/installation#next-steps","content":"Quick Start - Test your installation\nProvider Setup - Configure AI providers\nCLI Commands - Learn available commands\nExamples - See implementation patterns","hierarchy":{"lvl0":"Getting Started","lvl1":"Installation","lvl2":"Next Steps","lvl3":""}}, -{"objectID":"c44f69a5d19212c85a451486bef65999f3ef505311f283fa4a129d90603820ac","title":"⚙️ Provider Configuration Guide","url":"/docs/getting-started/provider-setup","content":"⚙️ Provider Configuration Guide\n\nNeuroLink supports multiple AI providers with flexible authentication methods. This guide covers complete setup for all supported providers.\n\nSupported Providers\n\nNeuroLink ships 40 providers in total. This guide walks through full environment-variable setup for the providers below; the complete roster — including the newer catalog providers and the embedding/media/decision-only providers — is indexed with setup guides at Provider Guides.\n\nProviders configured in this guide\nOpenAI - GPT-4o, GPT-4o-mini, GPT-4-turbo\nAmazon Bedrock - Claude 3.7 Sonnet, Claude 3.5 Sonnet, Claude 3 Haiku\nAmazon SageMaker - Custom models deployed on SageMaker endpoints\nGoogle Vertex AI - Gemini 3 Flash/Pro (preview), Gemini 2.5 Flash, Claude 4.0 Sonnet\nGoogle AI Studio - Gemini 1.5 Pro, Gemini 2.0 Flash, Gemini 1.5 Flash\nAnthropic - Claude 4.5 Opus/Sonnet/Haiku, Claude 4.0 Opus/Sonnet, Claude 3.7 Sonnet\nAzure OpenAI - GPT-4, GPT-3.5-Turbo\nLiteLLM - 100+ models from all providers via proxy server\nHugging Face - open models served by the unified router (Llama 3.x, Qwen 2.5, DeepSeek, Mistral)\nOllama - Local AI models including Llama 2, Code Llama, Mistral, Vicuna\nOpenRouter - 300+ models from every major lab via one aggregator endpoint\nMistral AI - Mistral Tiny, Small, Medium, and Large models\nDeepSeek - deepseek-chat (V3) and deepseek-reasoner (R1)\nNVIDIA NIM - Llama 3.3 70B and 400+ catalog models via NVIDIA hosted or self-hosted NIM\nLM Studio - Any model loaded in LM Studio desktop app (local, no API key required)\nllama.cpp - Any GGUF model served by llama-server (local, no API key required)\n\nOther providers (setup guides in the Provider Guides index)\n\nOnboarded via the zero-quirk OpenAI-wire-compatible catalog (Tier 2) — each has its own setup guide under :\nGroq - LPU-accelerated inference; default \nCerebras - Wafer-scale inference; default \nSambaNova - default \nTogether AI - default \nFireworks AI - default \nPerplexity - search-augmented models; default \nCloudflare Workers AI - edge inference\nxAI - Grok models; default \nBaseten - default ()\nGMI Cloud - default ()\nInception Labs - diffusion LLMs; default ()\nio.net Intelligence - decentralized GPU inference; default ()\nMancer - default (); no tool calling\nUpstage - Solar models; default ()\nAPI Route - OpenAI-compatible passthrough; default ()\n\nEmbedding, media-generation, and decision-only providers — not part of / provider selection in the same way, but each has a setup guide:\nCohere - chat + embeddings + reranking\nVoyage AI - embedding-only; default \nJina AI - embeddings + reranking; default \nReplicate, Stability AI, Ideogram, Recraft - direct image generation\nTypeSafe Jev - decision-only; serves , not /. Set (or for the gateway transport)\n\nVoice providers (TTS/STT/Realtime) are configured further down in this guide — see OpenAI TTS onward.\n\n💰 Model Availability & Cost Considerations\n\nImportant Notes:\nModel Availability: Specific models may not be available in all regions or require special access\nCost Variations: Pricing differs significantly between providers and models (e.g., Claude 3.5 Sonnet vs GPT-4o)\nRate Limits: Each provider has different rate limits and quota restrictions\nLocal vs Cloud: Ollama (local) has no per-request cost but requires hardware resources\nEnterprise Tiers: AWS Bedrock, Google Vertex AI, and Azure typically offer enterprise pricing\n\nBest Practices:\nUse with automatic provider selection for cost-optimized routing\nMonitor usage through built-in analytics to track costs\nConsider local models (Ollama) for development and testing\nCheck provider documentation for current pricing and availability\n\n🏢 Enterprise Proxy Support\n\nAll providers support corporate proxy environments automatically. Simply set environment variables:\n\nNo code changes required - NeuroLink automatically detects and uses proxy settings.\n\nFor detailed proxy setup → See Enterprise & Proxy Setup Guide\n\nOpenAI Configuration {#openai}\n\nBasic Setup\n\nOptional Configuration\n\nSupported Models\n(default) - Latest multimodal model\n- Cost-effective variant\n- High-performance model\n\nUsage Example\n\nTimeout Configuration\nDefault Timeout: 30 seconds\nSupported Formats: Milliseconds (), human-readable (, , )\nEnvironment Variable: (optional)\n\nAmazon Bedrock Configuration {#bedrock}\n\n🚨 Critical Setup Requirements\n\n⚠️ IMPORTANT: Anthropic Models Require Inference Profile ARN\n\nFor Anthropic Claude models in Bedrock, you MUST use the full inference profile ARN, not simple model names:\n\nBasic AWS Credentials\n\nSession Token Support (Development)\n\nFor temporary credentials (common in development environments):\n\nAvailable Inference Profile ARNs\n\nReplace with your AWS account ID:\n\nWhy Inference Profiles?\nCross-Region Access: Faster access across AWS regions\nBetter Performance: Optimized routing and response times\nHigher Availability: Improved model availability and reliability\nDifferent Permissions: Separate permission model from base models\n\nComplete Bedrock Configu","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"","lvl3":""}}, +{"objectID":"c44f69a5d19212c85a451486bef65999f3ef505311f283fa4a129d90603820ac","title":"⚙️ Provider Configuration Guide","url":"/docs/getting-started/provider-setup","content":"⚙️ Provider Configuration Guide\n\nNeuroLink supports multiple AI providers with flexible authentication methods. This guide covers complete setup for all supported providers.\n\nSupported Providers\n\nThis guide walks through full environment-variable setup for the providers below; the complete roster — including the newer catalog providers and the embedding/media/decision-only providers — is indexed with setup guides at Provider Guides.\n\nProviders configured in this guide\nOpenAI - GPT-4o, GPT-4o-mini, GPT-4-turbo\nAmazon Bedrock - Claude 3.7 Sonnet, Claude 3.5 Sonnet, Claude 3 Haiku\nAmazon SageMaker - Custom models deployed on SageMaker endpoints\nGoogle Vertex AI - Gemini 3 Flash/Pro (preview), Gemini 2.5 Flash, Claude 4.0 Sonnet\nGoogle AI Studio - Gemini 1.5 Pro, Gemini 2.0 Flash, Gemini 1.5 Flash\nAnthropic - Claude 4.5 Opus/Sonnet/Haiku, Claude 4.0 Opus/Sonnet, Claude 3.7 Sonnet\nAzure OpenAI - GPT-4, GPT-3.5-Turbo\nLiteLLM - 100+ models from all providers via proxy server\nHugging Face - open models served by the unified router (Llama 3.x, Qwen 2.5, DeepSeek, Mistral)\nOllama - Local AI models including Llama 2, Code Llama, Mistral, Vicuna\nOpenRouter - 300+ models from every major lab via one aggregator endpoint\nMistral AI - Mistral Tiny, Small, Medium, and Large models\nDeepSeek - deepseek-chat (V3) and deepseek-reasoner (R1)\nNVIDIA NIM - Llama 3.3 70B and 400+ catalog models via NVIDIA hosted or self-hosted NIM\nLM Studio - Any model loaded in LM Studio desktop app (local, no API key required)\nllama.cpp - Any GGUF model served by llama-server (local, no API key required)\n\nOther providers (setup guides in the Provider Guides index)\n\nOnboarded via the zero-quirk OpenAI-wire-compatible catalog (Tier 2) — each has its own setup guide under :\nGroq - LPU-accelerated inference; default \nCerebras - Wafer-scale inference; default \nSambaNova - default \nTogether AI - default \nFireworks AI - default \nPerplexity - search-augmented models; default \nCloudflare Workers AI - edge inference\nxAI - Grok models; default \nBaseten - default ()\nGMI Cloud - default ()\nInception Labs - diffusion LLMs; default ()\nio.net Intelligence - decentralized GPU inference; default ()\nMancer - default (); no tool calling\nUpstage - Solar models; default ()\nAPI Route - OpenAI-compatible passthrough; default ()\n\nEmbedding, media-generation, and decision-only providers — not part of / provider selection in the same way, but each has a setup guide:\nCohere - chat + embeddings + reranking\nVoyage AI - embedding-only; default \nJina AI - embeddings + reranking; default \nReplicate, Stability AI, Ideogram, Recraft - direct image generation\nTypeSafe Jev - decision-only; serves , not /. Set (or for the gateway transport)\nLaya - decision-only, open-weights; serves on a Laya server or LiteLLM proxy route you configure. Set + ; used when TypeSafe isn't configured\n\nVoice providers (TTS/STT/Realtime) are configured further down in this guide — see OpenAI TTS onward.\n\n💰 Model Availability & Cost Considerations\n\nImportant Notes:\nModel Availability: Specific models may not be available in all regions or require special access\nCost Variations: Pricing differs significantly between providers and models (e.g., Claude 3.5 Sonnet vs GPT-4o)\nRate Limits: Each provider has different rate limits and quota restrictions\nLocal vs Cloud: Ollama (local) has no per-request cost but requires hardware resources\nEnterprise Tiers: AWS Bedrock, Google Vertex AI, and Azure typically offer enterprise pricing\n\nBest Practices:\nUse with automatic provider selection for cost-optimized routing\nMonitor usage through built-in analytics to track costs\nConsider local models (Ollama) for development and testing\nCheck provider documentation for current pricing and availability\n\n🏢 Enterprise Proxy Support\n\nAll providers support corporate proxy environments automatically. Simply set environment variables:\n\nNo code changes required - NeuroLink automatically detects and uses proxy settings.\n\nFor detailed proxy setup → See Enterprise & Proxy Setup Guide\n\nOpenAI Configuration {#openai}\n\nBasic Setup\n\nOptional Configuration\n\nSupported Models\n(default) - Latest multimodal model\n- Cost-effective variant\n- High-performance model\n\nUsage Example\n\nTimeout Configuration\nDefault Timeout: 30 seconds\nSupported Formats: Milliseconds (), human-readable (, , )\nEnvironment Variable: (optional)\n\nAmazon Bedrock Configuration {#bedrock}\n\n🚨 Critical Setup Requirements\n\n⚠️ IMPORTANT: Anthropic Models Require Inference Profile ARN\n\nFor Anthropic Claude models in Bedrock, you MUST use the full inference profile ARN, not simple model names:\n\nBasic AWS Credentials\n\nSession Token Support (Development)\n\nFor temporary credentials (common in development environments):\n\nAvailable Inference Profile ARNs\n\nReplace with your AWS account ID:\n\nWhy Inference Profiles?\nCross-Region Access: Faster access across AWS regions\nBetter Performance: Optimized routing and response times\nHigher Availability: Improved model availability an","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"","lvl3":""}}, {"objectID":"44479af76ca558a6630aafa39578fb6543bd892844c3a6c5948ffd58a2d73287","title":"⚙️ Provider Configuration Guide","url":"/docs/getting-started/provider-setup#-provider-configuration-guide","content":"NeuroLink supports multiple AI providers with flexible authentication methods. This guide covers complete setup for all supported providers.","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"⚙️ Provider Configuration Guide","lvl3":""}}, -{"objectID":"ed4c54afa059a89c0cb4a7815f361a50933fe695ec62b3df5a0bd4ad2e25e431","title":"Supported Providers","url":"/docs/getting-started/provider-setup#supported-providers","content":"NeuroLink ships 40 providers in total. This guide walks through full environment-variable setup for the providers below; the complete roster — including the newer catalog providers and the embedding/media/decision-only providers — is indexed with setup guides at Provider Guides.","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"Supported Providers","lvl3":""}}, +{"objectID":"ed4c54afa059a89c0cb4a7815f361a50933fe695ec62b3df5a0bd4ad2e25e431","title":"Supported Providers","url":"/docs/getting-started/provider-setup#supported-providers","content":"This guide walks through full environment-variable setup for the providers below; the complete roster — including the newer catalog providers and the embedding/media/decision-only providers — is indexed with setup guides at Provider Guides.","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"Supported Providers","lvl3":""}}, {"objectID":"c0abb48e819713cb16c04d745137321d348118c823351fcf31ab014735dd036e","title":"Providers configured in this guide","url":"/docs/getting-started/provider-setup#providers-configured-in-this-guide","content":"OpenAI - GPT-4o, GPT-4o-mini, GPT-4-turbo\nAmazon Bedrock - Claude 3.7 Sonnet, Claude 3.5 Sonnet, Claude 3 Haiku\nAmazon SageMaker - Custom models deployed on SageMaker endpoints\nGoogle Vertex AI - Gemini 3 Flash/Pro (preview), Gemini 2.5 Flash, Claude 4.0 Sonnet\nGoogle AI Studio - Gemini 1.5 Pro, Gemini 2.0 Flash, Gemini 1.5 Flash\nAnthropic - Claude 4.5 Opus/Sonnet/Haiku, Claude 4.0 Opus/Sonnet, Claude 3.7 Sonnet\nAzure OpenAI - GPT-4, GPT-3.5-Turbo\nLiteLLM - 100+ models from all providers via proxy server\nHugging Face - open models served by the unified router (Llama 3.x, Qwen 2.5, DeepSeek, Mistral)\nOllama - Local AI models including Llama 2, Code Llama, Mistral, Vicuna\nOpenRouter - 300+ models from every major lab via one aggregator endpoint\nMistral AI - Mistral Tiny, Small, Medium, and Large models\nDeepSeek - deepseek-chat (V3) and deepseek-reasoner (R1)\nNVIDIA NIM - Llama 3.3 70B and 400+ catalog models via NVIDIA hosted or self-hosted NIM\nLM Studio - Any model loaded in LM Studio desktop app (local, no API key required)\nllama.cpp - Any GGUF model served by llama-server (local, no API key required)","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"Providers configured in this guide","lvl3":""}}, -{"objectID":"cc96a04c9f262c3c1f6f8b8e7c82407cc63406705250afd22947a3567e151721","title":"Other providers (setup guides in the Provider Guides index)","url":"/docs/getting-started/provider-setup#other-providers-setup-guides-in-the-provider-guides-index","content":"Onboarded via the zero-quirk OpenAI-wire-compatible catalog (Tier 2) — each has its own setup guide under :\nGroq - LPU-accelerated inference; default \nCerebras - Wafer-scale inference; default \nSambaNova - default \nTogether AI - default \nFireworks AI - default \nPerplexity - search-augmented models; default \nCloudflare Workers AI - edge inference\nxAI - Grok models; default \nBaseten - default ()\nGMI Cloud - default ()\nInception Labs - diffusion LLMs; default ()\nio.net Intelligence - decentralized GPU inference; default ()\nMancer - default (); no tool calling\nUpstage - Solar models; default ()\nAPI Route - OpenAI-compatible passthrough; default ()\n\nEmbedding, media-generation, and decision-only providers — not part of / provider selection in the same way, but each has a setup guide:\nCohere - chat + embeddings + reranking\nVoyage AI - embedding-only; default \nJina AI - embeddings + reranking; default \nReplicate, Stability AI, Ideogram, Recraft - direct image generation\nTypeSafe Jev - decision-only; serves , not /. Set (or for the gateway transport)\n\nVoice providers (TTS/STT/Realtime) are configured further down in this guide — see OpenAI TTS onward.","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"Other providers (setup guides in the Provider Guides index)","lvl3":""}}, +{"objectID":"cc96a04c9f262c3c1f6f8b8e7c82407cc63406705250afd22947a3567e151721","title":"Other providers (setup guides in the Provider Guides index)","url":"/docs/getting-started/provider-setup#other-providers-setup-guides-in-the-provider-guides-index","content":"Onboarded via the zero-quirk OpenAI-wire-compatible catalog (Tier 2) — each has its own setup guide under :\nGroq - LPU-accelerated inference; default \nCerebras - Wafer-scale inference; default \nSambaNova - default \nTogether AI - default \nFireworks AI - default \nPerplexity - search-augmented models; default \nCloudflare Workers AI - edge inference\nxAI - Grok models; default \nBaseten - default ()\nGMI Cloud - default ()\nInception Labs - diffusion LLMs; default ()\nio.net Intelligence - decentralized GPU inference; default ()\nMancer - default (); no tool calling\nUpstage - Solar models; default ()\nAPI Route - OpenAI-compatible passthrough; default ()\n\nEmbedding, media-generation, and decision-only providers — not part of / provider selection in the same way, but each has a setup guide:\nCohere - chat + embeddings + reranking\nVoyage AI - embedding-only; default \nJina AI - embeddings + reranking; default \nReplicate, Stability AI, Ideogram, Recraft - direct image generation\nTypeSafe Jev - decision-only; serves , not /. Set (or for the gateway transport)\nLaya - decision-only, open-weights; serves on a Laya server or LiteLLM proxy route you configure. Set + ; used when TypeSafe isn't configured\n\nVoice providers (TTS/STT/Realtime) are configured further down in this guide — see OpenAI TTS onward.","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"Other providers (setup guides in the Provider Guides index)","lvl3":""}}, {"objectID":"d17be3311722218280d0e76731e7344c620c18f5073e07abbc826f3922db976a","title":"💰 Model Availability & Cost Considerations","url":"/docs/getting-started/provider-setup#-model-availability-cost-considerations","content":"Important Notes:\nModel Availability: Specific models may not be available in all regions or require special access\nCost Variations: Pricing differs significantly between providers and models (e.g., Claude 3.5 Sonnet vs GPT-4o)\nRate Limits: Each provider has different rate limits and quota restrictions\nLocal vs Cloud: Ollama (local) has no per-request cost but requires hardware resources\nEnterprise Tiers: AWS Bedrock, Google Vertex AI, and Azure typically offer enterprise pricing\n\nBest Practices:\nUse with automatic provider selection for cost-optimized routing\nMonitor usage through built-in analytics to track costs\nConsider local models (Ollama) for development and testing\nCheck provider documentation for current pricing and availability","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"💰 Model Availability & Cost Considerations","lvl3":""}}, {"objectID":"3a1bd0ccb708d37a59e2db3650f689f77fba69e91ec88142c7c9b85f1c18057d","title":"🏢 Enterprise Proxy Support","url":"/docs/getting-started/provider-setup#-enterprise-proxy-support","content":"All providers support corporate proxy environments automatically. Simply set environment variables:\n\nNo code changes required - NeuroLink automatically detects and uses proxy settings.\n\nFor detailed proxy setup → See Enterprise & Proxy Setup Guide","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"🏢 Enterprise Proxy Support","lvl3":""}}, {"objectID":"1b0159447e4aee3ffd31efe28600cf755e4cc49d50f354e1a59d9574199213d2","title":"Supported Models","url":"/docs/getting-started/provider-setup#supported-models","content":"(default) - Latest multimodal model\n- Cost-effective variant\n- High-performance model","hierarchy":{"lvl0":"Getting Started","lvl1":"⚙️ Provider Configuration Guide","lvl2":"Supported Models","lvl3":""}}, @@ -5861,7 +5861,7 @@ {"objectID":"6ad36ba3488491bb913c7a77b86a278cc86509da88a054975bcded71f5c02130","title":"[OpenRouter](/docs/getting-started/providers/openrouter)","url":"/docs/getting-started/providers#openrouterdocsgetting-startedprovidersopenrouter","content":"300+ models from 60+ providers\n🌐 Single API for all major providers (Anthropic, OpenAI, Google, Meta, etc.)\n⚡ Automatic failover and routing\n💰 Competitive pricing with cost optimization\n🎯 Zero lock-in - switch models instantly\n📊 Usage tracking dashboard\n🆓 Free models available\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[OpenRouter](/docs/getting-started/providers/openrouter)","lvl3":""}}, {"objectID":"733a2fa8ea4c547cdcdf502f89f003a2208c0e5828e36e7faaaa94e0903d3d49","title":"[OpenAI Compatible](/docs/getting-started/providers/openai-compatible)","url":"/docs/getting-started/providers#openai-compatibledocsgetting-startedprovidersopenai-compatible","content":"OpenRouter, vLLM, LocalAI, and more\n🌐 100+ models through OpenRouter\n💻 Local deployment with vLLM\n🔓 Self-hosted with LocalAI\n🔄 Drop-in OpenAI replacement\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[OpenAI Compatible](/docs/getting-started/providers/openai-compatible)","lvl3":""}}, {"objectID":"fc415454bfde58f3634df544ef2adb88836c75fb1344b7042af7a9d6fdfa63d4","title":"[LiteLLM](/docs/getting-started/providers/litellm)","url":"/docs/getting-started/providers#litellmdocsgetting-startedproviderslitellm","content":"100+ providers through proxy\n🔄 Unified API for 100+ providers\n📊 Load balancing and fallbacks\n💰 Cost tracking\n🎯 Model routing\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[LiteLLM](/docs/getting-started/providers/litellm)","lvl3":""}}, -{"objectID":"faae7281e5a366b4892035d92012602a9e7aaa15cc9d96556b20493e25562498","title":"🧠 Decision-Only Providers {#decision-only-providers}","url":"/docs/getting-started/providers#-decision-only-providers-decision-only-providers","content":"The one provider that serves rather than /. It\nreturns typed, calibrated judgments and emits no text, so it never appears in\ngeneration fallback chains or the health sweep.","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"🧠 Decision-Only Providers {#decision-only-providers}","lvl3":""}}, +{"objectID":"faae7281e5a366b4892035d92012602a9e7aaa15cc9d96556b20493e25562498","title":"🧠 Decision-Only Providers {#decision-only-providers}","url":"/docs/getting-started/providers#-decision-only-providers-decision-only-providers","content":"The two providers that serve rather than /. Each\nreturns typed, calibrated judgments and emits no text, so neither appears in\ngeneration fallback chains or the health sweep.","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"🧠 Decision-Only Providers {#decision-only-providers}","lvl3":""}}, {"objectID":"50860d1f7a27dbe3ad7452ad7079ea437658db372b97081d0a509429369cacc6","title":"[TypeSafe (Jev)](/docs/getting-started/providers/typesafe)","url":"/docs/getting-started/providers#typesafe-jevdocsgetting-startedproviderstypesafe","content":"Typed, calibrated judgments instead of text\n🎯 / / answers, each with a calibrated confidence\n⚡ Latency flat in question count — 1 question ~393 ms, 400 questions ~465 ms\n💰 ~$0.00002 per decision (~$0.042/M input, output billed at zero)\n🔌 Two transports: TypeSafe direct, or the Vercel AI Gateway\n🛡️ Fails open — with no key configured, every consumer behaves exactly as before\n🔑 API key from console.typesafe.ai/keys\n🔄 Aliases: , \n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[TypeSafe (Jev)](/docs/getting-started/providers/typesafe)","lvl3":""}}, {"objectID":"43b838d60f2794ce46361196d6fc8a708b3a301398dd55f2c8015d567aa21c2c","title":"[Laya](/docs/getting-started/providers/laya)","url":"/docs/getting-started/providers#layadocsgetting-startedproviderslaya","content":"Open-weights decision provider — the same typed // answers as Jev, from Convai Innovations' Apache-2.0 checkpoints, at a Laya server or LiteLLM proxy route you configure\n🧭 Serves only; built-in features use it when its key and base URL are set and neither nor is\n📏 Refuses more than ~768 tokens of state (320 on ) before any network call, since its encoders read only 1,024 (512)\n🔌 No built-in endpoint: (or ) names a Laya server or a LiteLLM pass-through route to one\n🔑 is the key that endpoint accepts; on LiteLLM, the route must be in the key's Allowed Routes\n\nSetup Guide →","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"[Laya](/docs/getting-started/providers/laya)","lvl3":""}}, {"objectID":"c98403894a115164f29a019817b7f39c93b59623a9854b4575ad06bc5f597320","title":"🧩 Additional Catalog Providers","url":"/docs/getting-started/providers#-additional-catalog-providers","content":"Every provider below is a Tier-2 catalog entry — OpenAI-wire-compatible\nwith no behavioural quirks, so the whole integration is one JSON file under\n. Each page is generated from that file, which is\nalso what the CI onboarding gate reads.","hierarchy":{"lvl0":"Getting Started","lvl1":"AI Provider Guides","lvl2":"🧩 Additional Catalog Providers","lvl3":""}}, @@ -6550,10 +6550,10 @@ {"objectID":"14c8755cfa45311a2b8645d3729d760f513b4e768f4cbe6d75e1dcaaa68c7e94","title":"Configuration Reference","url":"/docs/getting-started/providers/together-ai#configuration-reference","content":"| Environment Variable | Required | Default |\n| -------------------- | -------- | ----------------------------------------- |\n| | Yes | — |\n| | No | |\n| | No | |","hierarchy":{"lvl0":"Getting Started","lvl1":"Together AI Provider Guide","lvl2":"Configuration Reference","lvl3":""}}, {"objectID":"73dacb114377509534778a0d5d1d386e64fe4e71f75f590903d6a5c3dda075c6","title":"Feature Support Matrix","url":"/docs/getting-started/providers/together-ai#feature-support-matrix","content":"| Feature | Support |\n| ----------------- | ----------------- |\n| Text generation | Yes |\n| Streaming | Yes |\n| Tool calling | Yes (model-dep.) |\n| Structured output | Yes (model-dep.) |\n| Vision | Yes (Llama 3.2 V) |\n| Embeddings | Limited |","hierarchy":{"lvl0":"Getting Started","lvl1":"Together AI Provider Guide","lvl2":"Feature Support Matrix","lvl3":""}}, {"objectID":"5a31d744d2d35af19cc87e27118e0d4057ab98d14fa7be6e9acb01dcbfc29f6a","title":"See Also","url":"/docs/getting-started/providers/together-ai#see-also","content":"Fireworks Provider\nGroq Provider","hierarchy":{"lvl0":"Getting Started","lvl1":"Together AI Provider Guide","lvl2":"See Also","lvl3":""}}, -{"objectID":"63e8301588e0fd846272e6b041d27fbe6a754543f38d949f07ce7e6f181a0ced","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe","content":"TypeSafe (Jev) Provider Guide\n\nThe only provider that serves rather than / — it\nreturns typed, calibrated judgments and emits no text at all.\n\nOverview\n\nTypeSafe's Jev is a \"System One\" model. You send one plus a map of\nnamed, typed questions; it returns one typed answer per question, all evaluated\nin a single parallel pass. Nothing has to be parsed back out of prose, and every\n/ answer carries a calibrated confidence rather than a\nself-reported one.\n\nBecause it emits no text, and are not available and\n throws — the same shape Voyage and Jina already use for\nembedding-only providers. Its descriptor declares ,\nwhich keeps it out of auto-select and the health sweep, so those throws are\nunreachable in normal use.\n\nThis is not , which\nscores an already-generated response with RAGAS scorers. Different feature,\ndifferent word.\n\nKey Facts\nProvider id: (aliases: , )\nInference kinds: only — the single provider of the 40 that does\nTool calling: none () — a decision model calls nothing\nHealth check: ; it is never probed with a live generation\nDefault decide timeout: 5000 ms ()\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.\n\nQuick Start\nGet an API key\n\nCreate one at console.typesafe.ai/keys.\nConfigure\nUse it\n\n returns on any failure. Use when you want the\nfailure to surface; it throws a whose carries a typed\n.\n\nFrom the CLI\n\n calls under the hood and reads\n from the environment like every other CLI command. See the\nCLI command reference.\n\nThe degradation contract\n\nSetting the key is the entire switch, and removing it is a complete undo.\nEvery internal consumer of fails open: with no decision provider\nconfigured, model routing, context budgeting, relevance compaction, tool routing\nand RAG planning all behave exactly as they did before. There is no\nconfiguration in which a missing, invalid, slow or unreachable decision model\nchanges NeuroLink's observable behaviour.\n\nA credential the service does not accept disables that provider instance rather\nthan paying a round trip on every later call to be told so again.\n\nTwo transports\n\nThe same model is reachable two ways, and the choice is made once in the\nconstructor.\n\n| | Direct | Vercel AI Gateway |\n| ------------------- | ------------------- | --------------------------------------------- |\n| Key | | |\n| Endpoint | | |\n| Endpoint override | | |\n| Model named in | request body | header |\n| Question vocabulary | | |\n| | on each answer | on |\n| Billed by | TypeSafe | Vercel |\n\nHolding both keys keeps the direct transport, so the confidence figures a\nhost already sees do not shift underneath it when a second key appears. Force\none with or\n. Either endpoint can be moved without a\nrelease: / for the direct\none, / for the\ngateway route.\n\n⚠️ The gateway refuses every request — free credits included — until the\nVercel team has a credit card on file, returning . That is an account state, not a bad key, and it\narrives before the model id is validated.\n\nFull detail, including the measured error table and why the distribution peak is\nnot a substitute for the reported confidence, is in\nThe inference type.\n\nWhat NeuroLink uses it for\n\n| Area | What the decision replaces |\n| ---------------------------------------------------------------- | --------------------------------------------------------------- |\n| Model routing | difficulty + capabilities + risk + model pick in one round trip |\n| Model catalogue | one over the registry ranks all N candidates at once |\n| Context budget | a rubric-placed scope reading lowers the compaction threshold |\n| Relevance compaction | per-message keep/drop, plus a gate on the generated summary |\n| Tool / MCP routing | one per server, replacing a 15s LLM call at ~400 ms |\n| RAG retrieval | per-query / hybrid / graph / rerank planning |\n\nLimits and gotchas\nTwo input ceilings, both enforced by the service: plus the longest\n single question ≈ 33,000 tokens, and plus all questions ≈\n 64,000. Exceeding either returns ","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"","lvl3":""}}, -{"objectID":"7320dc5177f0e320fe7887f77c01c06519bb6c0ddeb8a8b38ecfa013a9363a1f","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe#typesafe-jev-provider-guide","content":"The only provider that serves rather than / — it\nreturns typed, calibrated judgments and emits no text at all.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"TypeSafe (Jev) Provider Guide","lvl3":""}}, +{"objectID":"63e8301588e0fd846272e6b041d27fbe6a754543f38d949f07ce7e6f181a0ced","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe","content":"TypeSafe (Jev) Provider Guide\n\nOne of two providers that serve rather than / —\nit returns typed, calibrated judgments and emits no text at all. The other is\nLaya, an open-weights model you run yourself; TypeSafe runs first\nwhen both are configured.\n\nOverview\n\nTypeSafe's Jev is a \"System One\" model. You send one plus a map of\nnamed, typed questions; it returns one typed answer per question, all evaluated\nin a single parallel pass. Nothing has to be parsed back out of prose, and every\n/ answer carries a calibrated confidence rather than a\nself-reported one.\n\nBecause it emits no text, and are not available and\n throws — the same shape Voyage and Jina already use for\nembedding-only providers. Its descriptor declares ,\nwhich keeps it out of auto-select and the health sweep, so those throws are\nunreachable in normal use.\n\nThis is not , which\nscores an already-generated response with RAGAS scorers. Different feature,\ndifferent word.\n\nKey Facts\nProvider id: (aliases: , )\nInference kinds: only — one of two providers that do\n (the other is Laya); TypeSafe runs first when both are configured\nTool calling: none () — a decision model calls nothing\nHealth check: ; it is never probed with a live generation\nDefault decide timeout: 5000 ms ()\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.\n\nQuick Start\nGet an API key\n\nCreate one at console.typesafe.ai/keys.\nConfigure\nUse it\n\n returns on any failure. Use when you want the\nfailure to surface; it throws a whose carries a typed\n.\n\nFrom the CLI\n\n calls under the hood and reads\n from the environment like every other CLI command. See the\nCLI command reference.\n\nThe degradation contract\n\nSetting the key is the entire switch, and removing it is a complete undo.\nEvery internal consumer of fails open: with no decision provider\nconfigured, model routing, context budgeting, relevance compaction, tool routing\nand RAG planning all behave exactly as they did before. There is no\nconfiguration in which a missing, invalid, slow or unreachable decision model\nchanges NeuroLink's observable behaviour.\n\nA credential the service does not accept disables that provider instance rather\nthan paying a round trip on every later call to be told so again.\n\nTwo transports\n\nThe same model is reachable two ways, and the choice is made once in the\nconstructor.\n\n| | Direct | Vercel AI Gateway |\n| ------------------- | ------------------- | --------------------------------------------- |\n| Key | | |\n| Endpoint | | |\n| Endpoint override | | |\n| Model named in | request body | header |\n| Question vocabulary | | |\n| | on each answer | on |\n| Billed by | TypeSafe | Vercel |\n\nHolding both keys keeps the direct transport, so the confidence figures a\nhost already sees do not shift underneath it when a second key appears. Force\none with or\n. Either endpoint can be moved without a\nrelease: / for the direct\none, / for the\ngateway route.\n\n⚠️ The gateway refuses every request — free credits included — until the\nVercel team has a credit card on file, returning . That is an account state, not a bad key, and it\narrives before the model id is validated.\n\nFull detail, including the measured error table and why the distribution peak is\nnot a substitute for the reported confidence, is in\nThe inference type.\n\nWhat NeuroLink uses it for\n\n| Area | What the decision replaces |\n| ---------------------------------------------------------------- | --------------------------------------------------------------- |\n| Model routing | difficulty + capabilities + risk + model pick in one round trip |\n| Model catalogue | one over the registry ranks all N candidates at once |\n| Context budget | a rubric-placed scope reading lowers the compaction threshold |\n| Relevance compaction | per-message keep/drop, plus a gate on the generated summary |\n| Tool / MCP routing | one per server, replacing a 15s LLM call at ~400 ms |\n| RAG retrieval | per-query / hybrid / graph / rerank planning |\n\nLimits and gotchas\nT","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"","lvl3":""}}, +{"objectID":"7320dc5177f0e320fe7887f77c01c06519bb6c0ddeb8a8b38ecfa013a9363a1f","title":"TypeSafe (Jev) Provider Guide","url":"/docs/getting-started/providers/typesafe#typesafe-jev-provider-guide","content":"One of two providers that serve rather than / —\nit returns typed, calibrated judgments and emits no text at all. The other is\nLaya, an open-weights model you run yourself; TypeSafe runs first\nwhen both are configured.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"TypeSafe (Jev) Provider Guide","lvl3":""}}, {"objectID":"8cc640004e04379192695b8fcd3d49552ac0f4ebd763093bc610b4c93e823fee","title":"Overview","url":"/docs/getting-started/providers/typesafe#overview","content":"TypeSafe's Jev is a \"System One\" model. You send one plus a map of\nnamed, typed questions; it returns one typed answer per question, all evaluated\nin a single parallel pass. Nothing has to be parsed back out of prose, and every\n/ answer carries a calibrated confidence rather than a\nself-reported one.\n\nBecause it emits no text, and are not available and\n throws — the same shape Voyage and Jina already use for\nembedding-only providers. Its descriptor declares ,\nwhich keeps it out of auto-select and the health sweep, so those throws are\nunreachable in normal use.\n\nThis is not , which\nscores an already-generated response with RAGAS scorers. Different feature,\ndifferent word.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Overview","lvl3":""}}, -{"objectID":"4921fe7ae1c50c817be3b7742587c455084bee9d988015cf475edd5c3c196b84","title":"Key Facts","url":"/docs/getting-started/providers/typesafe#key-facts","content":"Provider id: (aliases: , )\nInference kinds: only — the single provider of the 40 that does\nTool calling: none () — a decision model calls nothing\nHealth check: ; it is never probed with a live generation\nDefault decide timeout: 5000 ms ()\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Key Facts","lvl3":""}}, +{"objectID":"4921fe7ae1c50c817be3b7742587c455084bee9d988015cf475edd5c3c196b84","title":"Key Facts","url":"/docs/getting-started/providers/typesafe#key-facts","content":"Provider id: (aliases: , )\nInference kinds: only — one of two providers that do\n (the other is Laya); TypeSafe runs first when both are configured\nTool calling: none () — a decision model calls nothing\nHealth check: ; it is never probed with a live generation\nDefault decide timeout: 5000 ms ()\nLatency: flat in question count — 1 question ~393 ms, 400 questions\n ~465 ms. Concurrent requests queue instead, so batch every question into one\n call rather than fanning out.\nCost: ~$0.042 per million input tokens, output billed at zero — about\n $0.00002 per decision. Output tokens are reported — measured 21 for a\n single question, converging to ~17.5 per question in a batch of eight — they\n are simply not charged.\nAccuracy is the trade: 67.8% on TypeSafe's own 711-case benchmark against\n Opus 5's 73.1%. Right for decisions that are gated and reversible; wrong for\n final answers.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"Key Facts","lvl3":""}}, {"objectID":"f5f7265f7ee9ba7ead6f6016a4b9232258ffaf4256e544874d6bcca8323dbb04","title":"1. Get an API key","url":"/docs/getting-started/providers/typesafe#1-get-an-api-key","content":"Create one at console.typesafe.ai/keys.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"1. Get an API key","lvl3":""}}, {"objectID":"a7cfa6599d9458e8c4a9bfcef1bba07a4948ebe6b0f30a1352df07a713040d8e","title":"3. Use it","url":"/docs/getting-started/providers/typesafe#3-use-it","content":"returns on any failure. Use when you want the\nfailure to surface; it throws a whose carries a typed\n.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"3. Use it","lvl3":""}}, {"objectID":"4ec7febf2956a104b0037472bab2e6c652e17df7f7cfdf69b73ebe2c5003f082","title":"From the CLI","url":"/docs/getting-started/providers/typesafe#from-the-cli","content":"calls under the hood and reads\n from the environment like every other CLI command. See the\nCLI command reference.","hierarchy":{"lvl0":"Getting Started","lvl1":"TypeSafe (Jev) Provider Guide","lvl2":"From the CLI","lvl3":""}}, @@ -7833,7 +7833,7 @@ {"objectID":"95f574ed3e7979eee5be188ccd16efb2639e30e8258e55366d7ccd344bbc3126","title":"Related Documentation","url":"/docs/implementation-guides/14-rag-document-processing#related-documentation","content":"Vector Store Integrations\nEvaluation and Scoring\nMaster Implementation Guide","hierarchy":{"lvl0":"Implementation Guides","lvl1":"RAG Document Processing - Implementation Guide","lvl2":"Related Documentation","lvl3":""}}, {"objectID":"60877eba17c1fe5c9fda2100a737f42fddb5c4c7083297e61a66884dbe5c326e","title":"NeuroLink","url":"/docs/","content":"🧠 NeuroLink\n The Pipe Layer of an AI Nervous System\n Provider Neurons for Every Major AI Vendor | 3 Inference Types (generate · stream · decide) | Voice (TTS/STT/Realtime) | 58+ MCP Tools | HITL Security | Redis Persistence\n\nNeuroLink is the pipe layer of an AI nervous system: one interface connecting provider neurons — every major AI vendor and local runtime — to the applications that consume them. Built-in tooling and an opinionated factory architecture mean adding a new provider, or a new capability, never touches application code. NeuroLink ships as both a TypeScript SDK and a professional CLI so teams can build, operate, and iterate on AI features quickly.\n\n🧠 What is NeuroLink?\n\nNeuroLink is the pipe layer of an AI nervous system. Providers — OpenAI, Anthropic, Google, AWS, Azure, DeepSeek, NVIDIA NIM, local runtimes like Ollama and llama.cpp, and dozens more — are the neurons: each generates a different kind of intelligence, at a different cost and latency. NeuroLink is the vascular layer that carries that intelligence, as a stream, to the applications that consume it, across three inference types: and produce text, produces a calibrated // judgment instead.\n\nExtracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — any provider you're building with, or any provider you add.\n\nWhy NeuroLink? Three genuine inference types, not one dressed up three ways — and produce text, while returns a typed, calibrated judgment ( / / ) with no text at all, for the routing and gating decisions the other two were never meant to make. Every neuron plugs into the same pipe. Switch providers with a single parameter change, leverage 64+ built-in tools and MCP servers, deploy with confidence using enterprise features like Redis memory and multi-provider failover, and optimize costs automatically with intelligent routing. Use it via our professional CLI or TypeScript SDK—whichever fits your workflow.\n\nWhere we're headed: We're building for the future of AI—edge-first execution and continuous streaming architectures that make AI practically free and universally available. Read our vision →\n\nGet Started in \\ Observability Guide\nServer Adapters -- Deploy NeuroLink as an HTTP API server with your framework of choice (Hono, Express, Fastify, Koa). Full CLI support with and commands for foreground/background modes, route management, and OpenAPI generation. -> Server Adapters Guide\nTitle Generation Events -- Emit real-time events when conversation titles are auto-generated. Listen to for session tracking. -> Conversation Memory Guide\nCustom Title Prompts -- Customize conversation title generation with environment variable. Use placeholder for dynamic prompts. -> Conversation Memory Guide\nVideo Generation -- Transform images into 8-second videos with synchronized audio using Google Veo 3.1 via Vertex AI. Supports 720p/1080p resolutions, portrait/landscape aspect ratios. -> Video Generation Guide\nImage Generation -- Generate images from text prompts using Gemini models via Vertex AI or Google AI Studio. Supports streaming mode with automatic file saving. -> Image Generation Guide\nHTTP/Streamable HTTP Transport for MCP -- Connect to remote MCP servers via HTTP with authentication headers, retry logic, and rate limiting. -> HTTP Transport Guide\nClaude Subscription (OAuth) Support -- Use your Claude Pro/Max/Team subscription with NeuroLink via OAuth authentication, no API key required. -> Subscription Guide\nGemini 3 Preview Support - Full support for gemini-3-flash-preview and gemini-3-pro-preview with extended thinking capabilities\nStructured Output with Zod Schemas -- Type-safe JSON generation with automatic validation using + in . -> Structured Output Guide\nCSV File Support -- Attach CSV files to prompts for AI-powered data analysis with auto-detection. -> CSV Guide\nPDF File Support -- Process PDF documents with native visual analysis for Vertex AI, Anthropic, Bedrock, AI Studio. -> PDF Guide\n50+ File Types -- Process Excel, Word, RTF, JSON, YAML, XML, HTML, SVG, Markdown, and 50+ code languages with intelligent content extraction. -> File Processors Guide\nLiteLLM Integration -- Access 100+ AI models from all major providers through unified interface. -> Setup Guide\nSageMaker Integration -- Deploy and use custom trained models on AWS infrastructure. -> Setup Guide\nOpenRouter Integration -- Access 300+ models from OpenAI, Anthropic, Google, Meta, and more through a single unified API. -> Setup Guide\nHuman-in-the-loop workflows -- Pause generation for user approval/input before tool execution. -> HITL Guide\nGuardrails middleware -- Block PII, profanity, and unsafe content with built-in filtering. -> Guardrails Guide\nContext summarization -- Automatic conversation compression for long-running sessions. -> Summarization Guide\nRedis conversation export -- Export ","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"","lvl3":""}}, {"objectID":"7bd402837d03dbfbb9586674bc914cfc6ee7691e33540eeabe6555363697375e","title":"🧠 What is NeuroLink?","url":"/docs/#-what-is-neurolink","content":"NeuroLink is the pipe layer of an AI nervous system. Providers — OpenAI, Anthropic, Google, AWS, Azure, DeepSeek, NVIDIA NIM, local runtimes like Ollama and llama.cpp, and dozens more — are the neurons: each generates a different kind of intelligence, at a different cost and latency. NeuroLink is the vascular layer that carries that intelligence, as a stream, to the applications that consume it, across three inference types: and produce text, produces a calibrated // judgment instead.\n\nExtracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — any provider you're building with, or any provider you add.\n\nWhy NeuroLink? Three genuine inference types, not one dressed up three ways — and produce text, while returns a typed, calibrated judgment ( / / ) with no text at all, for the routing and gating decisions the other two were never meant to make. Every neuron plugs into the same pipe. Switch providers with a single parameter change, leverage 64+ built-in tools and MCP servers, deploy with confidence using enterprise features like Redis memory and multi-provider failover, and optimize costs automatically with intelligent routing. Use it via our professional CLI or TypeScript SDK—whichever fits your workflow.\n\nWhere we're headed: We're building for the future of AI—edge-first execution and continuous streaming architectures that make AI practically free and universally available. Read our vision →\n\nGet Started in \\<5 Minutes →","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"🧠 What is NeuroLink?","lvl3":""}}, -{"objectID":"293082e7b04b9b0eb186f10fec40eefc1996d8eab9becff9e2b428731642291f","title":"What's New (Q1 2026)","url":"/docs/#whats-new-q1-2026","content":"| Feature | Version | Description | Guide |\n| ---------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- |\n| Inference Type | next | A third inference type alongside /: typed, calibrated // judgments in one parallel pass, ~400ms and ~$0.00002 per decision. First provider is TypeSafe Jev. Fail-open — a no-op without a key. | Decide Guide \\| TypeSafe Provider |\n| MCP Enhancements | v9.16.0 | Advanced MCP features: intelligent tool routing, result caching, request batching, tool annotations, elicitation protocol, custom server creation, multi-server management | MCP Enhancements Guide |\n| Context Compaction | v9.2.0 | 5-stage compaction pipeline (relevance, prune, deduplicate, summarize, truncate) with auto-detection, budget gate at 80% usage, per-provider token estimation | Context Compaction Guide |\n| File Processor System | v9.1.0 | 17 file processors across 6 categories with ProcessorRegistry, security sanitization, SVG text injection ","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"What's New (Q1 2026)","lvl3":""}}, +{"objectID":"293082e7b04b9b0eb186f10fec40eefc1996d8eab9becff9e2b428731642291f","title":"What's New (Q1 2026)","url":"/docs/#whats-new-q1-2026","content":"| Feature | Version | Description | Guide |\n| ---------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| Inference Type | next | A third inference type alongside /: typed, calibrated // judgments in one parallel pass, ~400ms and ~$0.00002 per decision. First provider is TypeSafe Jev; Laya, an open-weights model you run yourself, is the second. Fail-open — a no-op without a key. | Decide Guide \\| TypeSafe Provider \\| Laya Provider |\n| MCP Enhancements | v9.16.0 | Advanced MCP features: intelligent tool routing, result caching, request batching, tool annotations, elicitation protocol, custom server creation, multi-server management | MCP Enhancements Guide |\n| Context Compaction | v9.2.0 ","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"What's New (Q1 2026)","lvl3":""}}, {"objectID":"3e5ecc512d348470d92717ef33a2ff5067ffdd99c7ea212613df7986db133390","title":"Enterprise Security: Human-in-the-Loop (HITL)","url":"/docs/#enterprise-security-human-in-the-loop-hitl","content":"NeuroLink includes a HITL (Human-in-the-Loop) system for regulated industries and high-stakes AI operations:\n\n| Capability | Description | Use Case |\n| --------------------------- | ----------------------------------------------------------------------- | ------------------------------------------ |\n| Tool Approval Workflows | Require human approval before AI executes sensitive tools | Financial transactions, data modifications |\n| Output Validation | Route AI outputs through human review pipelines | Medical diagnosis, legal documents |\n| Confidence Thresholds | Automatically trigger human review below confidence level | Critical business decisions |\n| Complete Audit Trail | Audit logging to support your compliance program (HIPAA / SOC 2 / GDPR) | Regulated industries |\n\nEnterprise HITL Guide | Quick Start","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"Enterprise Security: Human-in-the-Loop (HITL)","lvl3":""}}, {"objectID":"c062d596f0f8940391d038563717712409b6055fdf70d92ab51cb35caf672663","title":"Get Started in Two Steps","url":"/docs/#get-started-in-two-steps","content":"`bash","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"Get Started in Two Steps","lvl3":""}}, {"objectID":"e9e7167316a23ea5fd5db23894b2e52405ba61f5ab57cc986d155490dd325674","title":"1. Run the interactive setup wizard (select providers, validate keys)","url":"/docs/#1-run-the-interactive-setup-wizard-select-providers-validate-keys","content":"pnpm dlx @juspay/neurolink setup","hierarchy":{"lvl0":"Docs","lvl1":"NeuroLink","lvl2":"1. Run the interactive setup wizard (select providers, validate keys)","lvl3":""}}, @@ -9121,10 +9121,10 @@ {"objectID":"4578b51cac195eb59c8ac256ef9f2491197666719997b358b74e2b57b5a414fb","title":"TTS / Realtime Errors","url":"/docs/reference/error-codes#tts-realtime-errors","content":"and carry provider-specific messages (e.g.\nsynthesis failure, WebSocket disconnect, function-call failure). Inspect\n for the underlying provider error.","hierarchy":{"lvl0":"Reference","lvl1":"Error Code Reference","lvl2":"TTS / Realtime Errors","lvl3":""}}, {"objectID":"3e6c6ffd80fa23ecf3686bacf0db97fb3e7e6f430e82b2523745c88e257b6c45","title":"Common Triggers","url":"/docs/reference/error-codes#common-triggers","content":"is thrown by \n when doesn't appear in the provider's\n list. The CLI infers the format from the\n file extension; the SDK requires you to pass it explicitly.\n Fix: either convert the audio to a supported format, or use a different\n STT provider. See for the\n Azure-MP3 case.\nis thrown when the buffer exceeds the per-call\n limit. Default is 25 MB (matches Whisper's documented\n ceiling). Override via .","hierarchy":{"lvl0":"Reference","lvl1":"Error Code Reference","lvl2":"Common Triggers","lvl3":""}}, {"objectID":"df1592c1b3ccdbed96fbbc998d2facc217e2652c2a3cbb2d630ed64f1c85fdb9","title":"Related Documentation","url":"/docs/reference/error-codes#related-documentation","content":"Troubleshooting Guide - Common issues and solutions\nConfiguration Reference - Environment variables and settings\nFAQ - Frequently asked questions\nProvider Feature Compatibility - Provider capabilities matrix","hierarchy":{"lvl0":"Reference","lvl1":"Error Code Reference","lvl2":"Related Documentation","lvl3":""}}, -{"objectID":"4dccb885a61a05030e1eeb0b2413af7062a282bcce434d742561026c5af1d015","title":"Frequently Asked Questions","url":"/docs/reference/faq","content":"Frequently Asked Questions\n\nCommon questions and answers about NeuroLink usage, configuration, and troubleshooting.\n\n🚀 Getting Started\n\nQ: What is NeuroLink?\n\nA: NeuroLink is an enterprise AI development platform that provides unified access to multiple AI providers (OpenAI, Google AI, Anthropic, AWS Bedrock, etc.) through a single SDK and CLI. It includes built-in tools, analytics, evaluation capabilities, and supports the Model Context Protocol (MCP) for extended functionality.\n\nQ: Which AI providers does NeuroLink support?\n\nA: NeuroLink ships 40 AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev (decision-only — serves , not /)\n\nSee Provider Setup for the complete roster with setup guides.\n\nVoice providers (a separate system from the 40 above):\nOpenAI TTS (TTS-1, TTS-1-HD, GPT-4o Audio)\nElevenLabs (Multilingual v2, Turbo v2.5, Flash v2.5)\nDeepgram (Nova-3, Nova-2, Enhanced — STT)\nAzure Speech (Azure Cognitive Services TTS + STT)\nGoogle TTS / STT (Google Cloud Speech)\nWhisper (OpenAI Whisper — STT)\nFish Audio (TTS)\nCartesia (TTS)\nOpenAI Realtime + Gemini Live (realtime voice APIs)\n\nMedia generation providers (image / video / music / avatar) — Kling, Runway, Replicate, Beatoven, Lyria, D-ID, HeyGen. See Media Generation for the full list.\n\nQ: Do I need to install anything?\n\nA: No installation required! You can use NeuroLink directly with :\n\nFor frequent use, you can install globally: \n\n🔧 Configuration\n\nQ: How do I set up API keys?\n\nA: Create a file in your project directory:\n\nNeuroLink automatically loads these environment variables.\n\nQ: Can I use NeuroLink behind a corporate proxy?\n\nA: Yes! NeuroLink automatically detects and uses corporate proxy settings:\n\nNo additional configuration needed.\n\nQ: How do I configure multiple environments (dev/staging/prod)?\n\nA: Use environment-specific files:\n\n🎯 Usage\n\nQ: What's the difference between CLI and SDK?\n\nA:\n\n| Feature | CLI | SDK |\n| -------------------- | ---------------------------- | ------------------------- |\n| Best for | Scripts, automation, testing | Applications, integration |\n| Installation | None required (npx) | npm install required |\n| Output | Text, JSON | Native JavaScript objects |\n| Batch processing | Built-in command | Manual implementation |\n| Learning curve | Low | Medium |\n\nQ: How do I choose the best provider for my use case?\n\nA: NeuroLink can auto-select the best provider, or you can choose based on:\nSpeed: Google AI (fastest responses)\nCoding: Anthropic Claude (best for code analysis)\nCreative: OpenAI (best for creative content)\nCost: Google AI Studio (free tier available)\nEnterprise: AWS Bedrock or Azure OpenAI\n\nQ: Can I use multiple providers in the same application?\n\nA: Yes! You can specify different providers for different requests:\n\n🔍 Troubleshooting\n\nQ: Why am I getting \"API key not found\" errors?\n\nA: Common solutions:\nCheck .env file exists and is in the correct directory\nVerify file format: No spaces around signs\nCheck file permissions: file should be readable\nVerify key format: Keys should start with provider-specific prefixes\n\nQ: Provider status shows \"Authentication failed\" - what should I do?\n\nA:\nVerify API key is correct and hasn't expired\nCheck account status - ensure billing is set up if required\nTest API key manually:\nCheck regional restrictions - some providers have geographic limitations\n\nQ: AWS Bedrock shows \"Not Authorized\" - how do I fix this?\n\nA: AWS Bedrock requires additional setup:\nRequest model access in AWS Bedrock console\nUse full inference profile ARN for Anthropic models:\nVerify IAM permissions include \nCheck AWS region - Bedrock isn't available in all regions\n\nQ: Google Vertex AI authenticati","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"","lvl3":""}}, +{"objectID":"4dccb885a61a05030e1eeb0b2413af7062a282bcce434d742561026c5af1d015","title":"Frequently Asked Questions","url":"/docs/reference/faq","content":"Frequently Asked Questions\n\nCommon questions and answers about NeuroLink usage, configuration, and troubleshooting.\n\n🚀 Getting Started\n\nQ: What is NeuroLink?\n\nA: NeuroLink is an enterprise AI development platform that provides unified access to multiple AI providers (OpenAI, Google AI, Anthropic, AWS Bedrock, etc.) through a single SDK and CLI. It includes built-in tools, analytics, evaluation capabilities, and supports the Model Context Protocol (MCP) for extended functionality.\n\nQ: Which AI providers does NeuroLink support?\n\nA: NeuroLink ships AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev, Laya (decision-only — serve , not /; TypeSafe first when both are configured)\n\nSee Provider Setup for the complete roster with setup guides.\n\nVoice providers (a separate system from the providers above):\nOpenAI TTS (TTS-1, TTS-1-HD, GPT-4o Audio)\nElevenLabs (Multilingual v2, Turbo v2.5, Flash v2.5)\nDeepgram (Nova-3, Nova-2, Enhanced — STT)\nAzure Speech (Azure Cognitive Services TTS + STT)\nGoogle TTS / STT (Google Cloud Speech)\nWhisper (OpenAI Whisper — STT)\nFish Audio (TTS)\nCartesia (TTS)\nOpenAI Realtime + Gemini Live (realtime voice APIs)\n\nMedia generation providers (image / video / music / avatar) — Kling, Runway, Replicate, Beatoven, Lyria, D-ID, HeyGen. See Media Generation for the full list.\n\nQ: Do I need to install anything?\n\nA: No installation required! You can use NeuroLink directly with :\n\nFor frequent use, you can install globally: \n\n🔧 Configuration\n\nQ: How do I set up API keys?\n\nA: Create a file in your project directory:\n\nNeuroLink automatically loads these environment variables.\n\nQ: Can I use NeuroLink behind a corporate proxy?\n\nA: Yes! NeuroLink automatically detects and uses corporate proxy settings:\n\nNo additional configuration needed.\n\nQ: How do I configure multiple environments (dev/staging/prod)?\n\nA: Use environment-specific files:\n\n🎯 Usage\n\nQ: What's the difference between CLI and SDK?\n\nA:\n\n| Feature | CLI | SDK |\n| -------------------- | ---------------------------- | ------------------------- |\n| Best for | Scripts, automation, testing | Applications, integration |\n| Installation | None required (npx) | npm install required |\n| Output | Text, JSON | Native JavaScript objects |\n| Batch processing | Built-in command | Manual implementation |\n| Learning curve | Low | Medium |\n\nQ: How do I choose the best provider for my use case?\n\nA: NeuroLink can auto-select the best provider, or you can choose based on:\nSpeed: Google AI (fastest responses)\nCoding: Anthropic Claude (best for code analysis)\nCreative: OpenAI (best for creative content)\nCost: Google AI Studio (free tier available)\nEnterprise: AWS Bedrock or Azure OpenAI\n\nQ: Can I use multiple providers in the same application?\n\nA: Yes! You can specify different providers for different requests:\n\n🔍 Troubleshooting\n\nQ: Why am I getting \"API key not found\" errors?\n\nA: Common solutions:\nCheck .env file exists and is in the correct directory\nVerify file format: No spaces around signs\nCheck file permissions: file should be readable\nVerify key format: Keys should start with provider-specific prefixes\n\nQ: Provider status shows \"Authentication failed\" - what should I do?\n\nA:\nVerify API key is correct and hasn't expired\nCheck account status - ensure billing is set up if required\nTest API key manually:\nCheck regional restrictions - some providers have geographic limitations\n\nQ: AWS Bedrock shows \"Not Authorized\" - how do I fix this?\n\nA: AWS Bedrock requires additional setup:\nRequest model access in AWS Bedrock console\nUse full inference profile ARN for Anthropic models:\nVerify IAM permissions include \nCheck AWS region - Bedrock isn't availabl","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"","lvl3":""}}, {"objectID":"430e71707f7db7f9866d7ab0b23f14f769c4812334161d853920bf4358ad8c17","title":"Frequently Asked Questions","url":"/docs/reference/faq#frequently-asked-questions","content":"Common questions and answers about NeuroLink usage, configuration, and troubleshooting.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Frequently Asked Questions","lvl3":""}}, {"objectID":"e9218a0cd45a1960da2da4d8d664f5a8b77494352527c7b4d1e6bc12c137400f","title":"Q: What is NeuroLink?","url":"/docs/reference/faq#q-what-is-neurolink","content":"A: NeuroLink is an enterprise AI development platform that provides unified access to multiple AI providers (OpenAI, Google AI, Anthropic, AWS Bedrock, etc.) through a single SDK and CLI. It includes built-in tools, analytics, evaluation capabilities, and supports the Model Context Protocol (MCP) for extended functionality.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: What is NeuroLink?","lvl3":""}}, -{"objectID":"c32f5558167b5817bde6577d51c68bbce793c3118b1192fdbf6c6e5958d3faff","title":"Q: Which AI providers does NeuroLink support?","url":"/docs/reference/faq#q-which-ai-providers-does-neurolink-support","content":"A: NeuroLink ships 40 AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev (decision-only — serves , not /)\n\nSee Provider Setup for the complete roster with setup guides.\n\nVoice providers (a separate system from the 40 above):\nOpenAI TTS (TTS-1, TTS-1-HD, GPT-4o Audio)\nElevenLabs (Multilingual v2, Turbo v2.5, Flash v2.5)\nDeepgram (Nova-3, Nova-2, Enhanced — STT)\nAzure Speech (Azure Cognitive Services TTS + STT)\nGoogle TTS / STT (Google Cloud Speech)\nWhisper (OpenAI Whisper — STT)\nFish Audio (TTS)\nCartesia (TTS)\nOpenAI Realtime + Gemini Live (realtime voice APIs)\n\nMedia generation providers (image / video / music / avatar) — Kling, Runway, Replicate, Beatoven, Lyria, D-ID, HeyGen. See Media Generation for the full list.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: Which AI providers does NeuroLink support?","lvl3":""}}, +{"objectID":"c32f5558167b5817bde6577d51c68bbce793c3118b1192fdbf6c6e5958d3faff","title":"Q: Which AI providers does NeuroLink support?","url":"/docs/reference/faq#q-which-ai-providers-does-neurolink-support","content":"A: NeuroLink ships AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include:\nOpenAI (GPT-4o, GPT-4.1, o3, o4-mini)\nGoogle AI Studio (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash)\nGoogle Vertex AI (Gemini 3, Claude via Vertex)\nAnthropic (Claude Opus 4.7, Sonnet 4.6, 4.5 Opus/Sonnet/Haiku)\nAWS Bedrock (Claude, Titan, Nova models)\nAzure OpenAI (GPT models)\nHugging Face (Open source models)\nOllama (Local AI models)\nMistral AI (Mistral models)\nLiteLLM (100+ models via proxy)\nAWS SageMaker (Custom endpoints)\nOpenAI-compatible (Any OpenAI-API-compatible endpoint)\nOpenRouter (300+ models via OpenRouter)\nDeepSeek (DeepSeek V3, R1)\nNVIDIA NIM (Llama 3.3 70B, 400+ catalog models)\nLM Studio (Local models loaded in LM Studio)\nllama.cpp (Local GGUF models via llama-server)\nGroq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route (zero-quirk OpenAI-wire-compatible catalog providers)\nCohere (chat, plus and reranking)\nVoyage AI, Jina AI (embedding and/or reranking only — no chat completions)\nTypeSafe Jev, Laya (decision-only — serve , not /; TypeSafe first when both are configured)\n\nSee Provider Setup for the complete roster with setup guides.\n\nVoice providers (a separate system from the providers above):\nOpenAI TTS (TTS-1, TTS-1-HD, GPT-4o Audio)\nElevenLabs (Multilingual v2, Turbo v2.5, Flash v2.5)\nDeepgram (Nova-3, Nova-2, Enhanced — STT)\nAzure Speech (Azure Cognitive Services TTS + STT)\nGoogle TTS / STT (Google Cloud Speech)\nWhisper (OpenAI Whisper — STT)\nFish Audio (TTS)\nCartesia (TTS)\nOpenAI Realtime + Gemini Live (realtime voice APIs)\n\nMedia generation providers (image / video / music / avatar) — Kling, Runway, Replicate, Beatoven, Lyria, D-ID, HeyGen. See Media Generation for the full list.","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: Which AI providers does NeuroLink support?","lvl3":""}}, {"objectID":"c7da7431322332ad0e9e6439acb89a8a9f6af428d15d84c94adeb886ea4fdaab","title":"Q: Do I need to install anything?","url":"/docs/reference/faq#q-do-i-need-to-install-anything","content":"A: No installation required! You can use NeuroLink directly with :\n\nFor frequent use, you can install globally:","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: Do I need to install anything?","lvl3":""}}, {"objectID":"379b5c908063649275668a696f71c0e87923dd0d011f221a19a43f69a483acea","title":"Q: How do I set up API keys?","url":"/docs/reference/faq#q-how-do-i-set-up-api-keys","content":"A: Create a file in your project directory:\n\n`bash","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":"Q: How do I set up API keys?","lvl3":""}}, {"objectID":"07c128fc6b5bcea9361ff3423a97e6a5a4e85fb14571f4a5d14bcb11eb389105","title":".env file","url":"/docs/reference/faq#env-file","content":"OPENAIAPIKEY=\"sk-your-openai-key\"\nGOOGLEAIAPI_KEY=\"AIza-your-google-ai-key\"\nANTHROPICAPIKEY=\"sk-ant-your-anthropic-key\"","hierarchy":{"lvl0":"Reference","lvl1":"Frequently Asked Questions","lvl2":".env file","lvl3":""}}, @@ -10246,8 +10246,8 @@ {"objectID":"1a0933649d2e2a9e068e520585c510992a787eb886f8661e6da36cf83ce7e3eb","title":"Multiple inputs","url":"/docs/skills/neurolink-guide/multimodal#multiple-inputs","content":"neurolink generate \"Compare\" --image ./a.png --image ./b.png --pdf ./docs.pdf\n`","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Multimodal Support","lvl2":"Multiple inputs","lvl3":""}}, {"objectID":"e2858a465b5d2bc8ef1db9834500597eb845bf3b40f3a0ba4b29e54033a4b23b","title":"File Size Considerations","url":"/docs/skills/neurolink-guide/multimodal#file-size-considerations","content":"| File Type | Recommended Max | Notes |\n| --------- | --------------- | -------------------- |\n| Images | 20MB | Resized if larger |\n| PDFs | 50MB | Page limit may apply |\n| CSV | 10MB | Use maxRows option |\n| Code | 100KB | Split large files |","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Multimodal Support","lvl2":"File Size Considerations","lvl3":""}}, {"objectID":"8d5a67779513f44569761e7ea2fa399de1aafd682a08e0b14c08850f0c223cb2","title":"Next Steps","url":"/docs/skills/neurolink-guide/multimodal#next-steps","content":"MCP tools - Add external tools\nRAG integration - Document-grounded generation\nProviders - Configure vision-capable providers","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Multimodal Support","lvl2":"Next Steps","lvl3":""}}, -{"objectID":"d9e19daaa21e9512ad71f1187055c474d4b51da0c78b56a197286c6403126b9a","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers","content":"NeuroLink Provider Configuration\n\nNeuroLink supports 40 AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev provider (serves , not /), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).\n\nCommon Providers\n\n| Provider | Enum Name | Aliases | Default Model |\n| ---------------- | -------------- | -------------- | --------------------------------------- |\n| OpenAI | | gpt, chatgpt | gpt-4o |\n| Anthropic | | claude | claude-3-5-sonnet-20241022 |\n| Google AI Studio | | gemini, google | gemini-2.5-flash |\n| Google Vertex AI | | google-vertex | gemini-2.5-flash |\n| AWS Bedrock | | aws-bedrock | anthropic.claude-3-sonnet-20240229-v1:0 |\n| Azure OpenAI | | azure | gpt-4o |\n| Mistral AI | | - | mistral-large |\n| Ollama | | - | llama3 |\n| LiteLLM | | - | varies |\n| AWS SageMaker | | - | custom |\n| Hugging Face | | hf | varies |\n| OpenRouter | | - | varies |\n| Gateway | | - | varies |\n\nOpenAI\n\nAvailable Models:\n- Latest GPT-4 Omni\n- Faster, cheaper\n- GPT-4 Turbo\n- Reasoning model\n- Smaller reasoning model\n\nAnthropic\n\nAvailable Models:\n- Latest Sonnet\n- Claude 3.7 Sonnet\n- Most capable\n- Fastest\n\nExtended Thinking:\n\nGoogle AI Studio\n\nAvailable Models:\n- Fast and capable\n- Most capable\n- Previous generation\n- Preview of Gemini 3\n\nGoogle Vertex AI\n\nAvailable Models:\n- Latest Gemini 3\n- Most capable Gemini 3\n- Fast\n- Previous gen capable\n\nExtended Thinking (Gemini 3):\n\nAWS Bedrock\n\nAvailable Models:\nAzure OpenAI\n\nMistral AI\n\nAvailable Models:\n- Most capable\n- Fast\n- Code specialized\n- Small\n\nOllama (Local)\n\nSetup:\n\nAvailable Models:\n- Meta Llama 3\n- Larger Llama 3\n- Mistral 7B\n- Code specialized\n- Microsoft Phi-3\n\nLiteLLM\n\nAWS SageMaker\n\nHugging Face\n\nOpenRouter\n\nProvider Fallback\n\nConfigure automatic fallback to another provider:\n\nCheck Provider Status\n\nProvider-Specific Options\n\nTemperature and Sampling\n\nSystem Prompts\n\nVision-Capable Models\n\nNot all models support image inputs:\n\n| Provider | Vision Models |\n| --------- | --------------------- |\n| OpenAI | gpt-4o, gpt-4-turbo |\n| Anthropic | All Claude 3 models |\n| Vertex | Gemini 2.5+, Gemini 3 |\n| Google AI | Gemini 2.5+, Gemini 3 |\n| Bedrock | Claude 3 models |\n\nNext Steps\nMultimodal inputs - Work with images and documents\nMCP tools - Add external tools\nRAG integration - Document-grounded generation","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"","lvl3":""}}, -{"objectID":"b66a3f5004ec2abbc42277973845a5485480e16a3ea54ea9f31e83978a9ee521","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers#neurolink-provider-configuration","content":"NeuroLink supports 40 AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev provider (serves , not /), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"NeuroLink Provider Configuration","lvl3":""}}, +{"objectID":"d9e19daaa21e9512ad71f1187055c474d4b51da0c78b56a197286c6403126b9a","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers","content":"NeuroLink Provider Configuration\n\nNeuroLink supports many AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev and Laya providers (serve , not /; TypeSafe first when both are configured), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).\n\nCommon Providers\n\n| Provider | Enum Name | Aliases | Default Model |\n| ---------------- | -------------- | -------------- | --------------------------------------- |\n| OpenAI | | gpt, chatgpt | gpt-4o |\n| Anthropic | | claude | claude-3-5-sonnet-20241022 |\n| Google AI Studio | | gemini, google | gemini-2.5-flash |\n| Google Vertex AI | | google-vertex | gemini-2.5-flash |\n| AWS Bedrock | | aws-bedrock | anthropic.claude-3-sonnet-20240229-v1:0 |\n| Azure OpenAI | | azure | gpt-4o |\n| Mistral AI | | - | mistral-large |\n| Ollama | | - | llama3 |\n| LiteLLM | | - | varies |\n| AWS SageMaker | | - | custom |\n| Hugging Face | | hf | varies |\n| OpenRouter | | - | varies |\n| Gateway | | - | varies |\n\nOpenAI\n\nAvailable Models:\n- Latest GPT-4 Omni\n- Faster, cheaper\n- GPT-4 Turbo\n- Reasoning model\n- Smaller reasoning model\n\nAnthropic\n\nAvailable Models:\n- Latest Sonnet\n- Claude 3.7 Sonnet\n- Most capable\n- Fastest\n\nExtended Thinking:\n\nGoogle AI Studio\n\nAvailable Models:\n- Fast and capable\n- Most capable\n- Previous generation\n- Preview of Gemini 3\n\nGoogle Vertex AI\n\nAvailable Models:\n- Latest Gemini 3\n- Most capable Gemini 3\n- Fast\n- Previous gen capable\n\nExtended Thinking (Gemini 3):\n\nAWS Bedrock\n\nAvailable Models:\nAzure OpenAI\n\nMistral AI\n\nAvailable Models:\n- Most capable\n- Fast\n- Code specialized\n- Small\n\nOllama (Local)\n\nSetup:\n\nAvailable Models:\n- Meta Llama 3\n- Larger Llama 3\n- Mistral 7B\n- Code specialized\n- Microsoft Phi-3\n\nLiteLLM\n\nAWS SageMaker\n\nHugging Face\n\nOpenRouter\n\nProvider Fallback\n\nConfigure automatic fallback to another provider:\n\nCheck Provider Status\n\nProvider-Specific Options\n\nTemperature and Sampling\n\nSystem Prompts\n\nVision-Capable Models\n\nNot all models support image inputs:\n\n| Provider | Vision Models |\n| --------- | --------------------- |\n| OpenAI | gpt-4o, gpt-4-turbo |\n| Anthropic | All Claude 3 models |\n| Vertex | Gemini 2.5+, Gemini 3 |\n| Google AI | Gemini 2.5+, Gemini 3 |\n| Bedrock | Claude 3 models |\n\nNext Steps\nMultimodal inputs - Work with images and documents\nMCP tools - Add external tools\nRAG integration - Document-grounded generation","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"","lvl3":""}}, +{"objectID":"b66a3f5004ec2abbc42277973845a5485480e16a3ea54ea9f31e83978a9ee521","title":"NeuroLink Provider Configuration","url":"/docs/skills/neurolink-guide/providers#neurolink-provider-configuration","content":"NeuroLink supports many AI providers through a unified API. This page highlights the most commonly-configured text providers — see the README provider table and the Provider Capabilities Audit for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev and Laya providers (serve , not /; TypeSafe first when both are configured), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live).","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"NeuroLink Provider Configuration","lvl3":""}}, {"objectID":"ca77e6781279a1e849f17c5edd76b2cc111546faa201202d3dc11820334170ce","title":"Common Providers","url":"/docs/skills/neurolink-guide/providers#common-providers","content":"| Provider | Enum Name | Aliases | Default Model |\n| ---------------- | -------------- | -------------- | --------------------------------------- |\n| OpenAI | | gpt, chatgpt | gpt-4o |\n| Anthropic | | claude | claude-3-5-sonnet-20241022 |\n| Google AI Studio | | gemini, google | gemini-2.5-flash |\n| Google Vertex AI | | google-vertex | gemini-2.5-flash |\n| AWS Bedrock | | aws-bedrock | anthropic.claude-3-sonnet-20240229-v1:0 |\n| Azure OpenAI | | azure | gpt-4o |\n| Mistral AI | | - | mistral-large |\n| Ollama | | - | llama3 |\n| LiteLLM | | - | varies |\n| AWS SageMaker | | - | custom |\n| Hugging Face | | hf | varies |\n| OpenRouter | | - | varies |\n| Gateway | | - | varies |","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"Common Providers","lvl3":""}}, {"objectID":"9c5370e9c625cd036ee1c3c8ba7c0bea9a5b6b68ca08ccd9c710d3e65429d843","title":"OpenAI","url":"/docs/skills/neurolink-guide/providers#openai","content":"`bash","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"OpenAI","lvl3":""}}, {"objectID":"fcc21bd062d4e7ab394bd0ad8a1257ec9ed367c205f68dab30479e6758c5cf20","title":"Environment","url":"/docs/skills/neurolink-guide/providers#environment","content":"OPENAIAPIKEY=sk-...\nOPENAIORGID=org-... # Optional\nOPENAIBASEURL=... # Optional, for proxies\ntypescript\nconst result = await neurolink.generate({\n input: { text: \"Hello\" },\n provider: \"openai\",\n model: \"gpt-4o\", // or gpt-4o-mini, gpt-4-turbo, o1, o1-mini\n});\ngpt-4ogpt-4o-minigpt-4-turboo1o1-mini` - Smaller reasoning model","hierarchy":{"lvl0":"Skills","lvl1":"NeuroLink Provider Configuration","lvl2":"Environment","lvl3":""}}, diff --git a/docs/about/nervous-system-model.md b/docs/about/nervous-system-model.md index c7a59ff03..412209ddd 100644 --- a/docs/about/nervous-system-model.md +++ b/docs/about/nervous-system-model.md @@ -12,7 +12,7 @@ NeuroLink is built around a biological metaphor — not as decoration, but as a ### Neurons — LLM Providers -Neurons are where intelligence is generated. In NeuroLink, neurons are the 40 AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar). +Neurons are where intelligence is generated. In NeuroLink, neurons are the AI providers, including: Anthropic, OpenAI, Google (AI Studio + Vertex), AWS (Bedrock + SageMaker), Azure, Mistral, LiteLLM, OpenRouter, Ollama, Hugging Face, DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, OpenAI-compatible endpoints, TypeSafe Jev and Laya (decision-only) — plus voice neurons (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Whisper, Deepgram, Azure STT, Google STT), realtime neurons (OpenAI Realtime, Gemini Live), and media-generation neurons (image, video, music, avatar). Each provider is a different type of neuron — different capabilities, different costs, different latency profiles. NeuroLink's ProviderRegistry gives you access to all of them through one interface, switchable with a single line. diff --git a/docs/features/index.md b/docs/features/index.md index b8428b475..0bf149729 100644 --- a/docs/features/index.md +++ b/docs/features/index.md @@ -12,37 +12,37 @@ Comprehensive guides for all NeuroLink features organized by category. Each guid ## Latest Features (Q1 2026) -| Feature | Description | -| ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| **[`decide` Inference Type](/docs/features/decide-inference-type)** | A third inference type alongside `generate`/`stream`: typed, calibrated `boolean`/`choice`/`score` judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision. First provider is [TypeSafe Jev](/docs/getting-started/providers/typesafe). Fail-open: a no-op without a key. | -| **[Model Routing with a Decision Model](/docs/features/classifier-router-jev-strategy)** | One round trip answers difficulty, capabilities, risk, context scope **and** the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). | -| **[Model Catalogue](/docs/features/classifier-router-catalog)** | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One `choice` question ranks all N candidates. | -| **[Context Budget](/docs/features/context-budget)** | A per-request compaction threshold derived from how much context the request actually needs. Only ever lowers the default, never raises it. | -| **[Relevance Compaction](/docs/features/relevance-compaction)** | Stage 0 of the compaction pipeline: asks which earlier messages the current request still needs, plus a quality gate on the generated summary. | -| **[Tool Routing with a Decision Model](/docs/features/tool-routing-decision-model)** | One calibrated yes/no per MCP server, dropping only on a confident "no" — replaces a 15s LLM call at ~400ms. | -| **[RAG Retrieval Planning](/docs/features/rag-retrieval-planning)** | Per-query `topK` breadth and whether to use hybrid / graph / rerank. Opt-in via `RAGPipeline`. | -| **[Real-time Voice Services](/docs/features/real-time-services)** | Bidirectional realtime voice APIs — OpenAI Realtime and Gemini Live. Full-duplex audio streaming with tool calls, barge-in, and interruption. | -| **[LiveKit Voice Agent](livekit-voice-agent.md)** | WebRTC voice agent using LiveKit for the real-time loop (transport, VAD, turn-taking, worker-per-call scaling) with NeuroLink as the brain (LLM, tools, memory). Cloud or self-hosted. | -| **[Provider Fallback & Model Chains](/docs/features/provider-fallback)** | `providerFallback` callback + `modelChain` config (v9.58.0) — centralized multi-provider fallback policy for resilient AI workflows. | -| **[Credential Validation](/docs/features/credential-validation)** | Pre-flight `sdk.checkCredentials()` API + typed `ModelAccessDeniedError` (v9.59.0) — actionable credential errors and validation before first call. | -| **[AutoResearch](autoresearch.md)** | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics, keeps improvements — runs unattended for hours. | -| **[MCP Enhancements](mcp-enhancements.md)** | Advanced MCP features: ToolRouter, ToolCache, RequestBatcher, tool annotations, elicitation protocol, and custom MCP server creation. | -| **[PPT Generation](ppt-generation.md)** | Generate professional PowerPoint presentations from text prompts with 35 slide types, 5 themes, and optional AI images. | -| **[Video Generation](video-generation.md)** | Generate videos from text prompts using RunwayML (ML5, ML6 Turbo models). | -| **[Image Generation with Gemini](../image-generation-streaming.md)** | Native image generation using Gemini 2.0 Flash Experimental with imagen-3.0-generate-002 model. | -| **[HTTP/Streamable HTTP Transport for MCP](../mcp-http-transport.md)** | Connect to remote MCP servers via HTTP with authentication, rate limiting, retry support, and session management. | -| **[Audio Input](audio-input.md)** | Real-time voice conversations with Gemini Live and audio streaming capabilities. | -| **[Server Adapters](../guides/server-adapters/index.md)** | Expose NeuroLink AI agents as HTTP APIs with Hono, Express, Fastify, and Koa. Production-ready with auth, rate limiting, and streaming. | -| **[RAG Document Processing](rag.md)** | Comprehensive document chunking (10 strategies), hybrid search (BM25 + vector), and reranking (5 types) for retrieval-augmented generation. | -| **[Context Compaction](context-compaction.md)** | 5-stage context compaction pipeline with automatic budget management, per-provider token estimation, and non-destructive message tagging. | -| **[Memory](memory.md)** | Per-user condensed memory that persists across conversations. LLM-powered condensation with S3, Redis, or SQLite storage backends. | -| **[Claude Subscription Support](claude-subscription.md)** | Multiple authentication methods for Claude (API key, OAuth) with support for Free, Pro, Max, and API tiers. | -| **[Client SDK](client-sdk.md)** | Type-safe HTTP, SSE, and WebSocket clients with React hooks and Vercel AI SDK adapter. | -| **[Proxy Accounts Endpoint](proxy-accounts-endpoint.md)** | GET /accounts: one row per account joining status, quota, and per-account token totals and API-equivalent cost. | -| **[Claude Proxy](claude-proxy.md)** | Multi-account Claude proxy with OAuth pooling, rate-limit failover, token refresh, and launchd daemon for crash recovery. | -| **[Proxy Peer Sharing](/docs/features/proxy-peer-sharing)** | Lend unused pool capacity to a peer's proxy with reserve floors, window slices, NeuroCoins, and instant pause or revoke. | -| **[Claude Proxy Observability](/docs/features/claude-proxy-observability)** | How to set up the local OpenObserve stack for the Claude proxy and read the dashboard for traffic health, failures, account routing, cache behavior, and trace drilldown. | -| **[Authentication Providers](/docs/features/authentication-providers)** | Secure AI endpoints with 11 auth providers (Auth0, Clerk, Firebase, Supabase, Cognito, Keycloak, WorkOS, Better Auth, OAuth2, JWT, Custom) with RBAC, session management, and rate limiting. | +| Feature | Description | +| ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **[`decide` Inference Type](/docs/features/decide-inference-type)** | A third inference type alongside `generate`/`stream`: typed, calibrated `boolean`/`choice`/`score` judgments in one parallel pass — no text. ~400ms, ~$0.00002 per decision. First provider is [TypeSafe Jev](/docs/getting-started/providers/typesafe); [Laya](/docs/getting-started/providers/laya), open-weights and self-hosted, is the second. Fail-open: a no-op without a key. | +| **[Model Routing with a Decision Model](/docs/features/classifier-router-jev-strategy)** | One round trip answers difficulty, capabilities, risk, context scope **and** the model pick, with asymmetric confidence bars (0.3 up / 0.6 down). | +| **[Model Catalogue](/docs/features/classifier-router-catalog)** | Ranks the 64-model registry as an addition to a host-declared pool, never a replacement. One `choice` question ranks all N candidates. | +| **[Context Budget](/docs/features/context-budget)** | A per-request compaction threshold derived from how much context the request actually needs. Only ever lowers the default, never raises it. | +| **[Relevance Compaction](/docs/features/relevance-compaction)** | Stage 0 of the compaction pipeline: asks which earlier messages the current request still needs, plus a quality gate on the generated summary. | +| **[Tool Routing with a Decision Model](/docs/features/tool-routing-decision-model)** | One calibrated yes/no per MCP server, dropping only on a confident "no" — replaces a 15s LLM call at ~400ms. | +| **[RAG Retrieval Planning](/docs/features/rag-retrieval-planning)** | Per-query `topK` breadth and whether to use hybrid / graph / rerank. Opt-in via `RAGPipeline`. | +| **[Real-time Voice Services](/docs/features/real-time-services)** | Bidirectional realtime voice APIs — OpenAI Realtime and Gemini Live. Full-duplex audio streaming with tool calls, barge-in, and interruption. | +| **[LiveKit Voice Agent](livekit-voice-agent.md)** | WebRTC voice agent using LiveKit for the real-time loop (transport, VAD, turn-taking, worker-per-call scaling) with NeuroLink as the brain (LLM, tools, memory). Cloud or self-hosted. | +| **[Provider Fallback & Model Chains](/docs/features/provider-fallback)** | `providerFallback` callback + `modelChain` config (v9.58.0) — centralized multi-provider fallback policy for resilient AI workflows. | +| **[Credential Validation](/docs/features/credential-validation)** | Pre-flight `sdk.checkCredentials()` API + typed `ModelAccessDeniedError` (v9.59.0) — actionable credential errors and validation before first call. | +| **[AutoResearch](autoresearch.md)** | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics, keeps improvements — runs unattended for hours. | +| **[MCP Enhancements](mcp-enhancements.md)** | Advanced MCP features: ToolRouter, ToolCache, RequestBatcher, tool annotations, elicitation protocol, and custom MCP server creation. | +| **[PPT Generation](ppt-generation.md)** | Generate professional PowerPoint presentations from text prompts with 35 slide types, 5 themes, and optional AI images. | +| **[Video Generation](video-generation.md)** | Generate videos from text prompts using RunwayML (ML5, ML6 Turbo models). | +| **[Image Generation with Gemini](../image-generation-streaming.md)** | Native image generation using Gemini 2.0 Flash Experimental with imagen-3.0-generate-002 model. | +| **[HTTP/Streamable HTTP Transport for MCP](../mcp-http-transport.md)** | Connect to remote MCP servers via HTTP with authentication, rate limiting, retry support, and session management. | +| **[Audio Input](audio-input.md)** | Real-time voice conversations with Gemini Live and audio streaming capabilities. | +| **[Server Adapters](../guides/server-adapters/index.md)** | Expose NeuroLink AI agents as HTTP APIs with Hono, Express, Fastify, and Koa. Production-ready with auth, rate limiting, and streaming. | +| **[RAG Document Processing](rag.md)** | Comprehensive document chunking (10 strategies), hybrid search (BM25 + vector), and reranking (5 types) for retrieval-augmented generation. | +| **[Context Compaction](context-compaction.md)** | 5-stage context compaction pipeline with automatic budget management, per-provider token estimation, and non-destructive message tagging. | +| **[Memory](memory.md)** | Per-user condensed memory that persists across conversations. LLM-powered condensation with S3, Redis, or SQLite storage backends. | +| **[Claude Subscription Support](claude-subscription.md)** | Multiple authentication methods for Claude (API key, OAuth) with support for Free, Pro, Max, and API tiers. | +| **[Client SDK](client-sdk.md)** | Type-safe HTTP, SSE, and WebSocket clients with React hooks and Vercel AI SDK adapter. | +| **[Proxy Accounts Endpoint](proxy-accounts-endpoint.md)** | GET /accounts: one row per account joining status, quota, and per-account token totals and API-equivalent cost. | +| **[Claude Proxy](claude-proxy.md)** | Multi-account Claude proxy with OAuth pooling, rate-limit failover, token refresh, and launchd daemon for crash recovery. | +| **[Proxy Peer Sharing](/docs/features/proxy-peer-sharing)** | Lend unused pool capacity to a peer's proxy with reserve floors, window slices, NeuroCoins, and instant pause or revoke. | +| **[Claude Proxy Observability](/docs/features/claude-proxy-observability)** | How to set up the local OpenObserve stack for the Claude proxy and read the dashboard for traffic health, failures, account routing, cache behavior, and trace drilldown. | +| **[Authentication Providers](/docs/features/authentication-providers)** | Secure AI endpoints with 11 auth providers (Auth0, Clerk, Firebase, Supabase, Cognito, Keycloak, WorkOS, Better Auth, OAuth2, JWT, Custom) with RBAC, session management, and rate limiting. | **Q1 2026 Highlights:** diff --git a/docs/getting-started/provider-setup.md b/docs/getting-started/provider-setup.md index 1315f655a..8913e1e30 100644 --- a/docs/getting-started/provider-setup.md +++ b/docs/getting-started/provider-setup.md @@ -4,7 +4,7 @@ NeuroLink supports multiple AI providers with flexible authentication methods. T ## Supported Providers -NeuroLink ships 40 providers in total. This guide walks through full environment-variable setup for the providers below; the complete roster — including the newer catalog providers and the embedding/media/decision-only providers — is indexed with setup guides at [Provider Guides](providers/index.md). +This guide walks through full environment-variable setup for the providers below; the complete roster — including the newer catalog providers and the embedding/media/decision-only providers — is indexed with setup guides at [Provider Guides](providers/index.md). ### Providers configured in this guide @@ -52,6 +52,7 @@ Embedding, media-generation, and decision-only providers — not part of `genera - **[Jina AI](providers/jina.md)** - embeddings + reranking; default `jina-embeddings-v3` - **[Replicate](providers/replicate.md)**, **[Stability AI](providers/stability.md)**, **[Ideogram](providers/ideogram.md)**, **[Recraft](providers/recraft.md)** - direct image generation - **[TypeSafe Jev](providers/typesafe.md)** - decision-only; serves `decide()`, not `generate()`/`stream()`. Set `TYPESAFE_API_KEY` (or `AI_GATEWAY_API_KEY` for the gateway transport) +- **[Laya](providers/laya.md)** - decision-only, open-weights; serves `decide()` on a Laya server or LiteLLM proxy route you configure. Set `LAYA_API_KEY` + `LAYA_BASE_URL`; used when TypeSafe isn't configured Voice providers (TTS/STT/Realtime) are configured further down in this guide — see [OpenAI TTS](#openai-tts) onward. diff --git a/docs/getting-started/providers/index.md b/docs/getting-started/providers/index.md index c09662412..2590240ed 100644 --- a/docs/getting-started/providers/index.md +++ b/docs/getting-started/providers/index.md @@ -454,8 +454,8 @@ Access multiple providers through unified interfaces: ## 🧠 Decision-Only Providers {#decision-only-providers} -The one provider that serves `decide` rather than `generate`/`stream`. It -returns typed, calibrated judgments and emits no text, so it never appears in +The two providers that serve `decide` rather than `generate`/`stream`. Each +returns typed, calibrated judgments and emits no text, so neither appears in generation fallback chains or the health sweep. ### [TypeSafe (Jev)](typesafe.md) diff --git a/docs/getting-started/providers/typesafe.md b/docs/getting-started/providers/typesafe.md index 9f8415bc9..2897e99b5 100644 --- a/docs/getting-started/providers/typesafe.md +++ b/docs/getting-started/providers/typesafe.md @@ -6,8 +6,10 @@ keywords: typesafe, jev, decide, decision model, calibrated confidence, routing, # TypeSafe (Jev) Provider Guide -**The only provider that serves `decide` rather than `generate`/`stream`** — it -returns typed, calibrated judgments and emits no text at all. +**One of two providers that serve `decide` rather than `generate`/`stream`** — +it returns typed, calibrated judgments and emits no text at all. The other is +[Laya](laya.md), an open-weights model you run yourself; TypeSafe runs first +when both are configured. --- @@ -32,7 +34,8 @@ unreachable in normal use. ### Key Facts - **Provider id**: `typesafe` (aliases: `jev`, `typesafe-ai`) -- **Inference kinds**: `decide` only — the single provider of the 40 that does +- **Inference kinds**: `decide` only — one of two providers that do + (the other is [Laya](laya.md)); TypeSafe runs first when both are configured - **Tool calling**: none (`toolSupport: "none"`) — a decision model calls nothing - **Health check**: `env-only`; it is never probed with a live generation - **Default decide timeout**: 5000 ms (`timeouts.decideMs`) diff --git a/docs/index.md b/docs/index.md index cdf40f4fb..011fbfcab 100644 --- a/docs/index.md +++ b/docs/index.md @@ -37,16 +37,16 @@ Extracted from production systems at Juspay, NeuroLink provides a practical, Typ ## What's New (Q1 2026) -| Feature | Version | Description | Guide | -| ---------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | -| **`decide` Inference Type** | next | A third inference type alongside `generate`/`stream`: typed, calibrated `boolean`/`choice`/`score` judgments in one parallel pass, ~400ms and ~$0.00002 per decision. First provider is TypeSafe Jev. Fail-open — a no-op without a key. | [Decide Guide](features/decide-inference-type.md) \| [TypeSafe Provider](getting-started/providers/typesafe.md) | -| **MCP Enhancements** | v9.16.0 | Advanced MCP features: intelligent tool routing, result caching, request batching, tool annotations, elicitation protocol, custom server creation, multi-server management | [MCP Enhancements Guide](features/mcp-enhancements.md) | -| **Context Compaction** | v9.2.0 | 5-stage compaction pipeline (relevance, prune, deduplicate, summarize, truncate) with auto-detection, budget gate at 80% usage, per-provider token estimation | [Context Compaction Guide](features/context-compaction.md) | -| **File Processor System** | v9.1.0 | 17 file processors across 6 categories with ProcessorRegistry, security sanitization, SVG text injection | [File Processors Guide](features/file-processors.md) | -| **Workflow Engine** | v8.42.0 | Multi-model orchestration with consensus, multi-judge, fallback, and adaptive workflows. Ensemble execution with intelligent scoring and evaluation. | [Workflow HLD](WORKFLOW-ENGINE-HLD.md) \| [Workflow LLD](WORKFLOW-ENGINE-LLD.md) | -| **Docusaurus Documentation** | v8.41.0 | Migrated from MkDocs to Docusaurus v3 with enhanced search, versioning, and modern UI. Automated doc syncing and LLM-friendly documentation. | [Documentation Site](https://docs.neurolink.ink) | -| **Image Generation with Gemini** | v8.31.0 | Native image generation using Gemini 2.0 Flash Experimental (`imagen-3.0-generate-002`). High-quality image synthesis directly from Google AI. | [Image Generation Guide](image-generation-streaming.md) | -| **HTTP/Streamable HTTP Transport** | v8.29.0 | Connect to remote MCP servers via HTTP with authentication headers, automatic retry with exponential backoff, and configurable rate limiting. | [HTTP Transport Guide](mcp-http-transport.md) | +| Feature | Version | Description | Guide | +| ---------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| **`decide` Inference Type** | next | A third inference type alongside `generate`/`stream`: typed, calibrated `boolean`/`choice`/`score` judgments in one parallel pass, ~400ms and ~$0.00002 per decision. First provider is TypeSafe Jev; [Laya](getting-started/providers/laya.md), an open-weights model you run yourself, is the second. Fail-open — a no-op without a key. | [Decide Guide](features/decide-inference-type.md) \| [TypeSafe Provider](getting-started/providers/typesafe.md) \| [Laya Provider](getting-started/providers/laya.md) | +| **MCP Enhancements** | v9.16.0 | Advanced MCP features: intelligent tool routing, result caching, request batching, tool annotations, elicitation protocol, custom server creation, multi-server management | [MCP Enhancements Guide](features/mcp-enhancements.md) | +| **Context Compaction** | v9.2.0 | 5-stage compaction pipeline (relevance, prune, deduplicate, summarize, truncate) with auto-detection, budget gate at 80% usage, per-provider token estimation | [Context Compaction Guide](features/context-compaction.md) | +| **File Processor System** | v9.1.0 | 17 file processors across 6 categories with ProcessorRegistry, security sanitization, SVG text injection | [File Processors Guide](features/file-processors.md) | +| **Workflow Engine** | v8.42.0 | Multi-model orchestration with consensus, multi-judge, fallback, and adaptive workflows. Ensemble execution with intelligent scoring and evaluation. | [Workflow HLD](WORKFLOW-ENGINE-HLD.md) \| [Workflow LLD](WORKFLOW-ENGINE-LLD.md) | +| **Docusaurus Documentation** | v8.41.0 | Migrated from MkDocs to Docusaurus v3 with enhanced search, versioning, and modern UI. Automated doc syncing and LLM-friendly documentation. | [Documentation Site](https://docs.neurolink.ink) | +| **Image Generation with Gemini** | v8.31.0 | Native image generation using Gemini 2.0 Flash Experimental (`imagen-3.0-generate-002`). High-quality image synthesis directly from Google AI. | [Image Generation Guide](image-generation-streaming.md) | +| **HTTP/Streamable HTTP Transport** | v8.29.0 | Connect to remote MCP servers via HTTP with authentication headers, automatic retry with exponential backoff, and configurable rate limiting. | [HTTP Transport Guide](mcp-http-transport.md) | - **External TracerProvider Support** -- Integrate NeuroLink with applications that already have OpenTelemetry instrumentation. Supports auto-detection and manual configuration. -> [Observability Guide](features/observability.md) - **Server Adapters** -- Deploy NeuroLink as an HTTP API server with your framework of choice (Hono, Express, Fastify, Koa). Full CLI support with `serve` and `server` commands for foreground/background modes, route management, and OpenAPI generation. -> [Server Adapters Guide](guides/server-adapters/index.md) @@ -194,7 +194,7 @@ NeuroLink is a comprehensive AI development platform. Every feature below is ava | **OpenAI Compatible** | Any OpenAI-compatible endpoint | Varies | ✅ Full | ✅ Production | [Setup Guide](getting-started/provider-setup.md#openai-compatible) | | **OpenRouter** | 200+ Models via OpenRouter | Varies | ✅ Full | ✅ Production | [Setup Guide](getting-started/providers/openrouter.md) | -This table highlights the most commonly used providers. NeuroLink also ships DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, xAI, Groq, Cerebras, SambaNova, Together AI, Fireworks, Perplexity, Cloudflare, Cohere, Voyage AI, Jina AI, Stability AI, Ideogram, Recraft, Replicate, plus TypeSafe Jev (a `decide()`-only provider for typed, calibrated decisions) and the full voice/media roster — see the [Provider Guides index](getting-started/providers/index.md) for all 40. +This table highlights the most commonly used providers. NeuroLink also ships DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, xAI, Groq, Cerebras, SambaNova, Together AI, Fireworks, Perplexity, Cloudflare, Cohere, Voyage AI, Jina AI, Stability AI, Ideogram, Recraft, Replicate, plus TypeSafe Jev and Laya (the two `decide()`-only providers for typed, calibrated decisions) and the full voice/media roster — see the [Provider Guides index](getting-started/providers/index.md) for the full roster. **[📖 Provider Comparison Guide](reference/provider-comparison.md)** - Detailed feature matrix and selection criteria **[🔬 Provider Feature Compatibility](reference/provider-feature-compatibility.md)** - Test-based compatibility reference for 19 features (dated snapshot covering a subset of the full provider list) diff --git a/docs/reference/faq.md b/docs/reference/faq.md index 4695ef0bd..8883d0dac 100644 --- a/docs/reference/faq.md +++ b/docs/reference/faq.md @@ -10,7 +10,7 @@ Common questions and answers about NeuroLink usage, configuration, and troublesh ### Q: Which AI providers does NeuroLink support? -**A:** NeuroLink ships 40 AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include: +**A:** NeuroLink ships AI providers for text generation, streaming, and decision-making — plus separate provider systems for voice and media generation. The text and multimodal providers include: - **OpenAI** (GPT-4o, GPT-4.1, o3, o4-mini) - **Google AI Studio** (Gemini 3 Flash/Pro, Gemini 2.5 Pro/Flash) @@ -32,11 +32,11 @@ Common questions and answers about NeuroLink usage, configuration, and troublesh - **Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage, API Route** (zero-quirk OpenAI-wire-compatible catalog providers) - **Cohere** (chat, plus `embed()` and reranking) - **Voyage AI**, **Jina AI** (embedding and/or reranking only — no chat completions) -- **TypeSafe Jev** (decision-only — serves `decide()`, not `generate()`/`stream()`) +- **TypeSafe Jev**, **Laya** (decision-only — serve `decide()`, not `generate()`/`stream()`; TypeSafe first when both are configured) See [Provider Setup](../getting-started/provider-setup.md) for the complete roster with setup guides. -Voice providers (a separate system from the 40 above): +Voice providers (a separate system from the providers above): - **OpenAI TTS** (TTS-1, TTS-1-HD, GPT-4o Audio) - **ElevenLabs** (Multilingual v2, Turbo v2.5, Flash v2.5) diff --git a/docs/skills/neurolink-guide/providers.md b/docs/skills/neurolink-guide/providers.md index 15741048e..09f456270 100644 --- a/docs/skills/neurolink-guide/providers.md +++ b/docs/skills/neurolink-guide/providers.md @@ -1,6 +1,6 @@ # NeuroLink Provider Configuration -NeuroLink supports 40 AI providers through a unified API. This page highlights the most commonly-configured text providers — see the [README provider table](https://github.com/juspay/neurolink/blob/main/README.md#supported-ai-providers) and the [Provider Capabilities Audit](https://github.com/juspay/neurolink/blob/main/docs/reference/provider-capabilities-audit.md) for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev provider (serves `decide()`, not `generate()`/`stream()`), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live). +NeuroLink supports many AI providers through a unified API. This page highlights the most commonly-configured text providers — see the [README provider table](https://github.com/juspay/neurolink/blob/main/README.md#supported-ai-providers) and the [Provider Capabilities Audit](https://github.com/juspay/neurolink/blob/main/docs/reference/provider-capabilities-audit.md) for the full matrix, including newer text providers (DeepSeek, NVIDIA NIM, LM Studio, llama.cpp, plus the Tier-2 catalog providers — Groq, Cerebras, SambaNova, Together AI, Fireworks AI, Perplexity, Cloudflare Workers AI, xAI, and more), the decision-only TypeSafe Jev and Laya providers (serve `decide()`, not `generate()`/`stream()`; TypeSafe first when both are configured), and voice providers (OpenAI TTS, ElevenLabs, Deepgram, Azure Speech, Whisper, Fish Audio, Cartesia, OpenAI Realtime, Gemini Live). ## Common Providers