Skip to content

feat(voice): voice/speech integration with TTS, STT, and realtime providers - #958

Closed
murdore wants to merge 1 commit into
releasefrom
feat/voice-speech-integration
Closed

murdore wants to merge 1 commit into
releasefrom
feat/voice-speech-integration

Conversation

@murdore

@murdore murdore commented Apr 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Adds voice/speech integration bridged into the existing TTSProcessor/generate() pipeline
  • 5 TTS providers, 4 STT providers, 2 Realtime providers registered with Processors
  • New STTProcessor mirroring TTSProcessor pattern with SpanType.STT observability

Consumer API

// TTS via generate (existing API, new providers)
await neurolink.generate({ input: { text: 'Hello' }, tts: { enabled: true, provider: 'elevenlabs' } });

// STT via generate (new)
await neurolink.generate({ input: { audio: buf }, stt: { enabled: true, provider: 'deepgram' } });

// Direct wrappers
await neurolink.synthesize('Hello', { provider: 'openai-tts' });
await neurolink.transcribe(audioBuffer, { provider: 'whisper' });

Code review fixes

  • GoogleTTS: implement getAccessToken() via google-auth-library
  • OpenAIRealtime: use model-provided call_id for function calls
  • stream-handler: drain pendingData on end()
  • AssemblyAI: WHATWG WebSocket compat (auth via query param)

Test plan

  • Type check, lint, build pass
  • Voice test suite 23/23
  • SDK generate+TTS end-to-end verified
  • CLI generate+stream with TTS flags verified
  • Live tests with ElevenLabs/Deepgram API keys

Summary by CodeRabbit

  • New Features

    • Added speech-to-text transcription with support for multiple providers (Whisper, Azure, Deepgram, Google, AssemblyAI, Gladia).
    • Added real-time voice conversation APIs for bidirectional audio interaction.
    • Expanded text-to-speech with new providers (OpenAI, ElevenLabs, Azure).
    • Added CLI support for voice synthesis provider selection and transcription.
    • Extended audio format support (m4a, flac, webm, mp4, mpeg, mpga).
  • Infrastructure

    • Made AWS SDK and Picovoice dependencies required.

Copilot AI review requested due to automatic review settings April 15, 2026 19:22
@vercel

vercel Bot commented Apr 15, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
neurolink Ready Ready Preview, Comment May 1, 2026 10:26am

@github-actions

github-actions Bot commented Apr 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Single Commit Policy - COMPLIANT

Status: Policy requirements met • 1 commit • Valid format • Ready for merge

📊 View validation details

📝 Commit Details

  • Hash: ad077b9677c61fcaa19c022aed2a24adf96365c8
  • Message: feat(voice): add multi-provider TTS, STT, and realtime voice integration
  • Author: Sachin Sharma

✅ Validation Results

  • Single commit requirement met
  • No merge commits in branch
  • Semantic commit message format verified
  • Ready for squash merge to release branch

🤖 Automated validation by NeuroLink Single Commit Enforcement

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@coderabbitai

coderabbitai Bot commented Apr 15, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@murdore has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 34 minutes and 46 seconds before requesting another review.

To keep reviews running without waiting, you can enable usage-based add-on for your organization. This allows additional reviews beyond the hourly cap. Account admins can enable it under billing.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: be2797bb-1e1a-451d-9a2d-3f32bd55bed0

📥 Commits

Reviewing files that changed from the base of the PR and between 4c93a47 and 1e4be5f.

⛔ Files ignored due to path filters (1)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (56)
  • .env.example
  • CHANGELOG.md
  • README.md
  • docs-site/scripts/sync-docs.ts
  • docs-site/sidebars.ts
  • docs/features/audio-input.md
  • docs/features/tts.md
  • docs/getting-started/environment-variables.md
  • docs/getting-started/provider-setup.md
  • docs/getting-started/providers/deepgram.md
  • docs/getting-started/providers/elevenlabs.md
  • docs/getting-started/providers/index.md
  • docs/getting-started/providers/openai-tts.md
  • docs/provider-integration/14-voice-speech-integration.md
  • docs/provider-integration/README.md
  • docs/reference/provider-comparison.md
  • docs/reference/provider-selection.md
  • examples/cli-examples.sh
  • memory-bank/voice-bridge-implementation-plan.md
  • memory-bank/voice-cleanup-plan.md
  • package.json
  • src/cli/factories/commandFactory.ts
  • src/cli/factories/sagemakerCommandFactory.ts
  • src/cli/loop/optionsSchema.ts
  • src/lib/core/baseProvider.ts
  • src/lib/factories/providerRegistry.ts
  • src/lib/neurolink.ts
  • src/lib/observability/exporters/laminarExporter.ts
  • src/lib/observability/exporters/posthogExporter.ts
  • src/lib/observability/utils/spanSerializer.ts
  • src/lib/server/voice/voiceWebSocketHandler.ts
  • src/lib/types/generate.ts
  • src/lib/types/index.ts
  • src/lib/types/realtime.ts
  • src/lib/types/server.ts
  • src/lib/types/span.ts
  • src/lib/types/stream.ts
  • src/lib/types/stt.ts
  • src/lib/types/tts.ts
  • src/lib/types/voice.ts
  • src/lib/utils/sttProcessor.ts
  • src/lib/voice/RealtimeVoiceAPI.ts
  • src/lib/voice/audio-utils.ts
  • src/lib/voice/errors.ts
  • src/lib/voice/index.ts
  • src/lib/voice/providers/AzureSTT.ts
  • src/lib/voice/providers/AzureTTS.ts
  • src/lib/voice/providers/DeepgramSTT.ts
  • src/lib/voice/providers/ElevenLabsTTS.ts
  • src/lib/voice/providers/GeminiLive.ts
  • src/lib/voice/providers/GoogleSTT.ts
  • src/lib/voice/providers/OpenAIRealtime.ts
  • src/lib/voice/providers/OpenAISTT.ts
  • src/lib/voice/providers/OpenAITTS.ts
  • src/lib/voice/stream-handler.ts
  • test/continuous-test-suite-voice.ts

Walkthrough

Added comprehensive speech-to-text, realtime voice, and new text-to-speech provider support. Introduced six STT providers (AssemblyAI, Azure, Deepgram, Gladia, Google, Whisper), two realtime providers (OpenAI, Gemini), four TTS providers, audio utilities for format detection and streaming, processor-based orchestration layers, and integrated STT into core generation workflows. Moved AWS SageMaker and Picovoice Cobra dependencies to required installation.

Changes

Cohort / File(s) Summary
STT Providers (Adapter Pattern)
src/lib/adapters/stt/assemblyaiSTTHandler.ts, azureSTTHandler.ts, deepgramSTTHandler.ts, gladiaSTTHandler.ts, googleSTTHandler.ts, whisperSTTHandler.ts
Six new STT provider handlers implementing batch transcription with audio upload/polling, validation, error handling, and result mapping. AssemblyAI, Azure, Deepgram, Gladia, Google, and Whisper (OpenAI). 600–700 LOC per handler.
STT Processor & Handler Layer
src/lib/voice/STTProvider.ts, src/lib/utils/sttProcessor.ts
Centralized STT processor orchestration with handler registry, provider validation, result enrichment with latency/provider metadata, and error wrapping. Span-based observability integration.
Realtime Voice Infrastructure
src/lib/voice/RealtimeVoiceAPI.ts
Core RealtimeProcessor (session/handler registry management) and BaseRealtimeHandler (shared base class with event emission, session creation, state management).
Realtime Providers
src/lib/voice/providers/GeminiLive.ts, OpenAIRealtime.ts
WebSocket-based realtime voice implementations for Gemini and OpenAI with audio/text I/O, function call handling, stream parsing, and connection lifecycle.
TTS Providers
src/lib/voice/providers/AzureTTS.ts, ElevenLabsTTS.ts, GoogleTTS.ts, OpenAITTS.ts
Four new text-to-speech handlers (Azure, ElevenLabs, Google, OpenAI) with voice enumeration, synthesis, SSML/format handling, and voice caching.
STT Providers (New Handler Pattern)
src/lib/voice/providers/AzureSTT.ts, DeepgramSTT.ts, GoogleSTT.ts, OpenAISTT.ts
Refactored STT handlers conforming to STTHandler interface with streaming support, configuration validation, and buffer-based streaming fallback for non-streaming providers.
Audio Utilities & Streaming
src/lib/voice/audio-utils.ts, stream-handler.ts
Format detection, MIME mapping, duration calculation (WAV/MP3/Opus), PCM encoding/resampling, WAV file construction, chunk-based audio streaming with backpressure, stream merging/splitting utilities.
Voice Error Handling
src/lib/voice/errors.ts
Comprehensive error hierarchy (VoiceError, STTError, RealtimeError) with factory methods for provider-scoped error codes, categories, severities, and retriability flags.
Voice Module Exports
src/lib/voice/index.ts
Consolidated entrypoint re-exporting error classes, processors (STT, Realtime), audio utilities, stream handlers, and provider implementations (TTS/STT/Realtime) with *Handler aliases.
Type Definitions
src/lib/types/stt.ts, realtime.ts, voice.ts, tts.ts (updated), generate.ts (updated), stream.ts (updated), span.ts (updated)
STT option/result/segment/language/error types; Realtime session/config/handler/event types; Voice provider/capability abstractions; AudioFormat expansion; Generation/Stream option extensions with STT audio input; SpanType.STT addition.
Core Integration
src/lib/core/baseProvider.ts
STT transcription hook in generate() flow: validates provider config, transcribes audio when options.stt.enabled, sets transcription result as prompt, attaches result to final output; TTS provider selection preference for explicit provider override.
NeuroLink Public APIs
src/lib/neurolink.ts
Added transcribe(), synthesize(), startRealtimeVoice() public methods with processor delegation and provider defaults (whisper, google-ai, openai-realtime).
Provider Registry
src/lib/factories/providerRegistry.ts
Dynamic registration of new TTS handlers (OpenAI, ElevenLabs, Azure) and STT handlers (Whisper, Deepgram, Google, Azure) plus Realtime handlers (OpenAI, Gemini) with try/catch suppression for missing dependencies.
CLI Enhancements
src/cli/factories/commandFactory.ts
Added STT CLI flags (--stt, --stt-provider, --stt-language, --input-audio), TTS provider override flag (--tts-provider), and payload wiring for both generate and stream paths.
SageMaker & Cobra Integration
src/cli/factories/sagemakerCommandFactory.ts, src/lib/server/voice/voiceWebSocketHandler.ts
Replaced dynamic imports with static imports for @aws-sdk/client-sagemaker and @picovoice/cobra-node; synchronous initialization and validation.
Observability
src/lib/observability/exporters/laminarExporter.ts, posthogExporter.ts, utils/spanSerializer.ts
Added SpanType.STT mapping to exporters (Laminar: "custom", PostHog: "ai_stt_transcription", LangSmith: "chain").
Test Infrastructure
test/continuous-test-suite-voice.ts
New test runner validating voice module imports, processor/handler exports, provider availability with credential checks, audio-utils/stream-handler modules, and error reporting.
Dependencies & Metadata
package.json, CHANGELOG.md, memory-bank/voice-bridge-implementation-plan.md
Version downgrade 9.55.4→9.55.2; moved @aws-sdk/client-sagemaker and @picovoice/cobra-node from optionalDependencies to dependencies; added voice implementation plan documentation.
Type Schema Updates
src/cli/loop/optionsSchema.ts, src/lib/types/server.ts
Excluded stt from CLI primitive schema; removed unused CobraInstance type export.

Estimated code review effort

🎯 5 (Critical) | ⏱️ ~120 minutes

Rationale: Heterogeneous changes across 50+ files with high-density logic in provider implementations (6 STT adapters + 4 STT handlers + 4 TTS handlers + 2 realtime handlers = 16 complex providers), 500–700 LOC per provider. Core generation flow modifications, processor-based orchestration patterns, WebSocket/async streaming integration, comprehensive type system expansion (3 new type modules with 1900+ LOC), error hierarchy with factory patterns, and observability instrumentation. Diverse patterns (REST HTTP, WebSocket, async generators, event emitters) require separate reasoning per provider and module. Dependency reshuffling and CLI integration add complexity.

Possibly related PRs

Suggested labels

released

Suggested reviewers

  • vigneshJuspay

Poem

🐰 Hark! The whisper transforms to text,
While voices real-time now connect,
From Deepgram, Azure, Gemini's gleam—
Speech flows through processors supreme! 🎙️✨

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/voice-speech-integration

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds a new voice/speech subsystem to the NeuroLink SDK + CLI, covering TTS, STT, and realtime voice providers, with supporting registries/factories and a continuous test script.

Changes:

  • Introduces voice infrastructure (registry/factory, composite orchestration, realtime base processor, stream utilities, error types).
  • Adds provider implementations for TTS (Google/OpenAI/ElevenLabs/Azure), STT (Whisper/OpenAI, Deepgram, Google, Azure), and realtime (OpenAI Realtime, Gemini Live).
  • Extends the CLI with neurolink voice subcommands and adds a continuous voice integration test script.

Reviewed changes

Copilot reviewed 33 out of 33 changed files in this pull request and generated 8 comments.

Show a summary per file
File Description
test/continuous-test-suite-voice.ts Adds a runnable continuous “smoke suite” for provider/module presence and basic wiring.
src/lib/voice/voiceRegistry.ts Implements provider registry (IDs, aliases, metadata, type filtering).
src/lib/voice/voiceAgent.ts Adds a high-level voice-to-voice agent (STT → LLM → TTS) + realtime session support.
src/lib/voice/stream-handler.ts Adds chunking/backpressure helpers and async-iterable adapters for audio streams.
src/lib/voice/providers/OpenAITTS.ts Implements OpenAI TTS via REST API.
src/lib/voice/providers/OpenAISTT.ts Implements OpenAI Whisper STT via multipart/form-data REST API.
src/lib/voice/providers/OpenAIRealtime.ts Implements OpenAI Realtime bidirectional voice over WebSocket.
src/lib/voice/providers/GoogleTTS.ts Implements Google Cloud TTS via REST API + voice listing cache.
src/lib/voice/providers/GoogleSTT.ts Implements Google Cloud STT via REST API + placeholder streaming.
src/lib/voice/providers/GeminiLive.ts Implements Gemini Live bidirectional voice over WebSocket.
src/lib/voice/providers/ElevenLabsTTS.ts Implements ElevenLabs TTS via REST API + voice listing cache.
src/lib/voice/providers/DeepgramSTT.ts Implements Deepgram STT (batch + WebSocket streaming).
src/lib/voice/providers/AzureTTS.ts Implements Azure TTS via REST API + voice listing cache.
src/lib/voice/providers/AzureSTT.ts Implements Azure STT via REST API + placeholder streaming.
src/lib/voice/index.ts Exports the voice module public surface area.
src/lib/voice/errors.ts Adds shared error classes/codes for voice, STT, realtime.
src/lib/voice/compositeVoice.ts Adds a TTS+STT orchestrator with history and conversation-turn helper.
src/lib/voice/audio-utils.ts Adds audio format detection, duration estimation, and WAV/PCM utilities.
src/lib/voice/STTProvider.ts Adds a centralized STT processor (similar to existing TTSProcessor pattern).
src/lib/voice/RealtimeVoiceAPI.ts Adds centralized realtime processor + base handler abstraction.
src/lib/types/tts.ts Expands AudioFormat union to cover more formats used by STT APIs.
src/lib/types/index.ts Exports the consolidated voice types.
src/lib/adapters/stt/whisperSTTHandler.ts Adds an adapter-style Whisper STT handler (non-voice-module path).
src/cli/parser.ts Wires the new voice command group into the CLI.
src/cli/commands/voice.ts Implements `neurolink voice synthesize

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/lib/voice/providers/GoogleTTS.ts Outdated
Comment on lines +90 to +376
isConfigured(): boolean {
return this.apiKey !== null || this.credentialsPath !== null;
}

async getVoices(languageCode?: string): Promise<TTSVoice[]> {
if (!this.isConfigured()) {
throw new TTSError({
code: TTS_ERROR_CODES.PROVIDER_NOT_CONFIGURED,
message: "Google TTS is not configured",
category: ErrorCategory.CONFIGURATION,
severity: ErrorSeverity.HIGH,
retriable: false,
});
}

// Return cached voices if valid and no language filter
if (
this.voicesCache &&
Date.now() - this.voicesCache.timestamp < GoogleTTS.CACHE_TTL_MS &&
!languageCode
) {
return this.voicesCache.voices;
}

try {
const params = new URLSearchParams();
if (languageCode) {
params.set("languageCode", languageCode);
}

const url = this.apiKey
? `${this.baseUrl}/voices?key=${this.apiKey}&${params.toString()}`
: `${this.baseUrl}/voices?${params.toString()}`;

const response = await fetch(url, {
method: "GET",
headers: {
...(this.credentialsPath && !this.apiKey
? { Authorization: `Bearer ${await this.getAccessToken()}` }
: {}),
},
});

if (!response.ok) {
throw new Error(`HTTP ${response.status}`);
}

const data = (await response.json()) as GoogleListVoicesResponse;

const voices: TTSVoice[] = data.voices.map((voice) => ({
id: voice.name,
name: voice.name,
languageCode: voice.languageCodes[0] ?? "en-US",
languageCodes: voice.languageCodes,
gender: this.mapGender(voice.ssmlGender),
type: this.extractVoiceType(voice.name),
naturalSampleRateHertz: voice.naturalSampleRateHertz,
}));

// Cache if no language filter
if (!languageCode) {
this.voicesCache = { voices, timestamp: Date.now() };
}

return voices;
} catch (err: unknown) {
const errorMessage =
err instanceof Error ? err.message : String(err || "Unknown error");
logger.error(`[GoogleTTSHandler] Failed to get voices: ${errorMessage}`);
throw new TTSError({
code: TTS_ERROR_CODES.SYNTHESIS_FAILED,
message: `Failed to get voices: ${errorMessage}`,
category: ErrorCategory.NETWORK,
severity: ErrorSeverity.MEDIUM,
retriable: true,
originalError: err instanceof Error ? err : undefined,
});
}
}

async synthesize(text: string, options: TTSOptions = {}): Promise<TTSResult> {
if (!this.isConfigured()) {
throw new TTSError({
code: TTS_ERROR_CODES.PROVIDER_NOT_CONFIGURED,
message: "Google TTS is not configured",
category: ErrorCategory.CONFIGURATION,
severity: ErrorSeverity.HIGH,
retriable: false,
});
}

const startTime = Date.now();
const googleOptions = options as GoogleTTSOptions;

try {
// Detect if text is SSML
const isSSML = text.trim().startsWith("<speak");

// Build synthesis input
const input: GoogleSynthesisInput = isSSML ? { ssml: text } : { text };

// Parse voice and language from voice name or use defaults
const voiceName = options.voice ?? "en-US-Neural2-C";
const languageCode = this.extractLanguageCode(voiceName);

// Build voice selection
const voice: GoogleVoiceSelectionParams = {
languageCode,
name: voiceName,
};

// Build audio config
const audioConfig: GoogleAudioConfig = {
audioEncoding: this.getEncoding(options.format ?? "mp3"),
speakingRate: options.speed ?? 1.0,
pitch: options.pitch ?? 0.0,
volumeGainDb: options.volumeGainDb ?? 0.0,
};

if (googleOptions.sampleRateHertz) {
audioConfig.sampleRateHertz = googleOptions.sampleRateHertz;
}

if (googleOptions.effectsProfileId) {
audioConfig.effectsProfileId = googleOptions.effectsProfileId;
}

// Build request
const request: GoogleSynthesizeRequest = {
input,
voice,
audioConfig,
};

const url = this.apiKey
? `${this.baseUrl}/text:synthesize?key=${this.apiKey}`
: `${this.baseUrl}/text:synthesize`;

const response = await fetch(url, {
method: "POST",
headers: {
"Content-Type": "application/json",
...(this.credentialsPath && !this.apiKey
? { Authorization: `Bearer ${await this.getAccessToken()}` }
: {}),
},
body: JSON.stringify(request),
});

if (!response.ok) {
const errorData = await response
.json()
.catch(() => Object.create(null) as Record<string, unknown>);
const errorMessage =
(errorData as { error?: { message?: string } }).error?.message ||
`HTTP ${response.status}`;
throw new Error(errorMessage);
}

const data = (await response.json()) as GoogleSynthesizeResponse;
const latency = Date.now() - startTime;

// Decode base64 audio
const audioBuffer = Buffer.from(data.audioContent, "base64");

const result: TTSResult = {
buffer: audioBuffer,
format: options.format ?? "mp3",
size: audioBuffer.length,
voice: voiceName,
sampleRate:
googleOptions.sampleRateHertz ??
this.getDefaultSampleRate(options.format),
metadata: {
latency,
provider: "google-tts",
encoding: audioConfig.audioEncoding,
},
};

logger.info(
`[GoogleTTSHandler] Synthesized ${audioBuffer.length} bytes in ${latency}ms`,
);

return result;
} catch (err: unknown) {
if (err instanceof TTSError) {
throw err;
}

const errorMessage =
err instanceof Error ? err.message : String(err || "Unknown error");
logger.error(`[GoogleTTSHandler] Synthesis failed: ${errorMessage}`);
throw new TTSError({
code: TTS_ERROR_CODES.SYNTHESIS_FAILED,
message: `Synthesis failed: ${errorMessage}`,
category: ErrorCategory.EXECUTION,
severity: ErrorSeverity.HIGH,
retriable: true,
context: { textLength: text.length },
originalError: err instanceof Error ? err : undefined,
});
}
}

/**
* Map SSML gender to standard gender type
*/
private mapGender(ssmlGender: string): "male" | "female" | "neutral" {
switch (ssmlGender?.toUpperCase()) {
case "MALE":
return "male";
case "FEMALE":
return "female";
default:
return "neutral";
}
}

/**
* Extract voice type from voice name
*/
private extractVoiceType(
name: string,
): "standard" | "wavenet" | "neural" | "chirp" | "unknown" {
const nameLower = name.toLowerCase();
if (nameLower.includes("neural2") || nameLower.includes("neural")) {
return "neural";
}
if (nameLower.includes("wavenet")) {
return "wavenet";
}
if (nameLower.includes("standard")) {
return "standard";
}
if (nameLower.includes("chirp")) {
return "chirp";
}
return "unknown";
}

/**
* Extract language code from voice name
*/
private extractLanguageCode(voiceName: string): string {
// Voice names are formatted like "en-US-Neural2-C"
const match = voiceName.match(/^([a-z]{2}-[A-Z]{2})/);
return match ? match[1] : "en-US";
}

/**
* Get encoding string for audio format
*/
private getEncoding(format: AudioFormat): string {
const encodings: Partial<Record<AudioFormat, string>> = {
mp3: "MP3",
wav: "LINEAR16",
ogg: "OGG_OPUS",
opus: "OGG_OPUS",
};
return encodings[format] ?? "MP3";
}

/**
* Get default sample rate for format
*/
private getDefaultSampleRate(format?: AudioFormat): number {
switch (format) {
case "wav":
return 16000;
case "ogg":
case "opus":
return 24000;
default:
return 24000;
}
}

/**
* Get access token from service account (placeholder)
*/
private async getAccessToken(): Promise<string> {
logger.warn(
"[GoogleTTSHandler] Service account auth not implemented, use API key",
);
return "";
}

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

isConfigured() returns true when GOOGLE_APPLICATION_CREDENTIALS is set, but getAccessToken() is a placeholder that returns an empty string. In that configuration the provider will attempt authenticated requests with Authorization: Bearer and fail at runtime. Either implement service-account token acquisition (e.g., via google-auth-library) or treat "credentialsPath only" as not configured until token support is implemented.

Copilot uses AI. Check for mistakes.
Comment on lines +504 to +510
JSON.stringify({
type: "conversation.item.create",
item: {
type: "function_call_output",
call_id: name, // Note: This should be the actual call_id from the event
output: JSON.stringify(result),
},

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

call_id is being set to the function name, not the model-provided call identifier. OpenAI Realtime expects the original call_id from the function-call event; using the name will cause tool outputs to be ignored or misrouted. Capture call_id from the incoming event (e.g., include it in the parsed event shape) and echo that value here.

Copilot uses AI. Check for mistakes.
Comment thread src/lib/voice/voiceAgent.ts Outdated
Comment on lines +338 to +346
const sessionConfig: RealtimeConfig = {
provider: "openai",
voice: this.config.voiceSettings?.voiceId,
instructions: this.config.systemPrompt,
temperature: 0.7,
turnDetection: "server_vad",
...this.config.realtimeConfig,
...config,
};

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RealtimeConfig consumers in this PR use config.systemPrompt (e.g., the provider session setup), but startRealtimeSession() sets instructions instead. This means the system prompt likely won’t be applied for realtime sessions unless providers also read instructions. Set systemPrompt (or set both fields consistently) when building sessionConfig.

Copilot uses AI. Check for mistakes.
Comment thread src/lib/voice/compositeVoice.ts Outdated
Comment on lines +217 to +230
// Convert to voice module STT options format
const voiceSTTOptions: STTOptions = {
language: options.language,
format: options.format as STTOptions["format"],
diarization: options.diarization,
punctuate: options.punctuate,
wordTimestamps: options.wordTimestamps,
confidenceThreshold: (options as { confidenceThreshold?: number })
.confidenceThreshold,
};

const mergedOptions: STTOptions = {
...this.config.defaultSTTOptions,
...voiceSTTOptions,

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

transcribe() rebuilds a new STTOptions object with only a small subset of fields, which will silently drop valid options like sampleRate, punctuation, speakerDiarization, profanityFilter, etc. Since options is already Partial<STTOptions>, merge it directly with defaultSTTOptions (and only normalize aliases if needed) to avoid losing settings.

Suggested change
// Convert to voice module STT options format
const voiceSTTOptions: STTOptions = {
language: options.language,
format: options.format as STTOptions["format"],
diarization: options.diarization,
punctuate: options.punctuate,
wordTimestamps: options.wordTimestamps,
confidenceThreshold: (options as { confidenceThreshold?: number })
.confidenceThreshold,
};
const mergedOptions: STTOptions = {
...this.config.defaultSTTOptions,
...voiceSTTOptions,
// Merge caller-provided STT options directly so valid fields are not dropped.
// Preserve the existing normalization for fields that require explicit typing.
const mergedOptions: STTOptions = {
...this.config.defaultSTTOptions,
...options,
format:
(options.format as STTOptions["format"] | undefined) ??
this.config.defaultSTTOptions?.format,
confidenceThreshold:
(options as { confidenceThreshold?: number }).confidenceThreshold ??
this.config.defaultSTTOptions?.confidenceThreshold,

Copilot uses AI. Check for mistakes.
Comment on lines +145 to +170
end(): void {
if (this.isEnded) {
return;
}

this.isEnded = true;
this.clearBufferTimeout();

// Emit final chunk with remaining data
if (this.buffer.length > 0) {
const durationMs =
(this.buffer.length /
this.config.bytesPerSample /
this.config.sampleRate) *
1000;

const chunk: AudioStreamChunk = {
data: this.buffer,
index: this.chunkIndex++,
isFinal: true,
format: this.config.format,
sampleRate: this.config.sampleRate,
timestampMs: this.timestampMs,
durationMs,
};

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If backpressure occurred, data is buffered in pendingData, but end() only flushes this.buffer and ignores pendingData. That can drop audio data if the caller ends the stream while paused. Consider concatenating/processing pendingData before emitting the final chunk (or ensure end() drains pending data first).

Copilot uses AI. Check for mistakes.
Comment thread src/cli/commands/voice.ts Outdated
Comment on lines +194 to +205
// Detect format from extension
const ext = path.extname(inputPath).toLowerCase().slice(1);
const format = argv.format ?? ext;

spinner.text = `Transcribing with ${providerName}...`;

const options: STTOptions = {
language: argv.language,
format: format as STTOptions["format"],
diarization: argv.diarization,
wordTimestamps: argv.wordTimestamps,
};

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The CLI accepts/auto-detects formats like flac, m4a, webm, etc., but there’s no validation against the selected STT provider’s getSupportedFormats(). This can lead to confusing runtime failures (or wrong MIME types) when a provider doesn’t support the detected format. Validate the resolved format against provider.getSupportedFormats() and fail fast with a helpful message listing supported formats.

Copilot uses AI. Check for mistakes.
Comment thread test/continuous-test-suite-voice.ts Outdated
Comment on lines +243 to +246
const realtimeProviders = [
{ name: "OpenAIRealtime", envKey: "OPENAI_API_KEY" },
{ name: "GeminiLive", envKey: "GOOGLE_AI_STUDIO_API_KEY" },
];

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GeminiLive reads its API key from process.env.GOOGLE_API_KEY, but this test suite checks GOOGLE_AI_STUDIO_API_KEY. That mismatch will incorrectly report missing credentials (or skip unexpectedly). Update the env var used here to match the provider implementation (or support both).

Copilot uses AI. Check for mistakes.
Comment on lines +8 to +10
import * as fs from "fs";
import * as path from "path";

Copilot AI Apr 15, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fs and path are imported but never used in this script, which will trigger lint/tsc unused import warnings. Remove them or use them for any planned file IO to keep the test runner clean.

Copilot uses AI. Check for mistakes.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Note

Due to the large number of review comments, Critical severity comments were prioritized as inline comments.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/lib/neurolink.ts (1)

11666-11782: ⚠️ Potential issue | 🟠 Major

The lazy-init path creates an unusable CompositeVoice.

On a fresh NeuroLink instance these methods call initializeVoice() with {}, which creates a CompositeVoice with no ttsProvider / sttProvider. The very next call then fails inside CompositeVoice.synthesize() / CompositeVoice.transcribe() with provider-not-configured.

🔧 Proposed fix
   async synthesize(text: string, options?: TTSOptions): Promise<TTSResult> {
-    // Initialize voice if not already done
-    if (!this.compositeVoice) {
-      await this.initializeVoice();
-    }
-
     if (!this.compositeVoice) {
       throw new VoiceError({
         code: VOICE_ERROR_CODES.INVALID_CONFIGURATION,
         message: "Voice not initialized. Call initializeVoice() first.",
@@
   async transcribe(
     audio: Buffer | ArrayBuffer,
     options?: STTOptions,
   ): Promise<STTResult> {
-    // Initialize voice if not already done
-    if (!this.compositeVoice) {
-      await this.initializeVoice();
-    }
-
     if (!this.compositeVoice) {
       throw new VoiceError({
         code: VOICE_ERROR_CODES.INVALID_CONFIGURATION,
         message: "Voice not initialized. Call initializeVoice() first.",

Also applies to: 11859-11878

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 11666 - 11782, The lazy-init path can
create a CompositeVoice with no providers, causing
CompositeVoice.synthesize()/transcribe() to fail; fix by adding explicit
provider checks after initialization: in synthesize (and similarly in
transcribe) verify this.compositeVoice and that a TTS/STT provider is configured
(e.g., this.compositeVoice.ttsProvider / .sttProvider or a hasTTS()/hasSTT()
helper) and throw a descriptive VoiceError (use
VOICE_ERROR_CODES.INVALID_CONFIGURATION) that instructs callers to call
initializeVoice(...) with a ttsProvider/sttProvider or configure SDK defaults;
alternatively, update initializeVoice to pull default providers from NeuroLink
configuration when config is empty so CompositeVoice is created with usable
providers. Ensure references: initializeVoice, synthesize, transcribe,
CompositeVoice, VOICE_ERROR_CODES, VoiceError.
🟠 Major comments (25)
test/continuous-test-suite-voice.ts-303-310 (1)

303-310: ⚠️ Potential issue | 🟠 Major

Make missing audio-utils exports fail this suite.

exists || true forces the check to pass every time, so a missing export only shows up as skipped noise instead of a real failure. This turns the export verification into a no-op.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-voice.ts` around lines 303 - 310, The test loop is
forcing every export check to pass by using `exists || true`; change the pass
condition to use the actual `exists` boolean so missing exports fail the suite.
Locate the loop that iterates `for (const fn of expectedFunctions)` and update
the `recordTest` call for `audio-utils.${fn} exists` so the second argument is
`exists` (not `exists || true`) and keep the failure message (`exists ?
undefined : "Function not found"`) as-is; this ensures `audioUtils` missing
exports cause a real test failure.
src/lib/voice/providers/AzureTTS.ts-164-196 (1)

164-196: ⚠️ Potential issue | 🟠 Major

Reject unsupported Azure output formats instead of reporting them as generated.

mapFormat() only supports four formats, but the result object still reports options.format at Line 196. If a caller asks for one of the newly added formats, Azure will produce mp3/opus while the SDK labels the file as the requested format.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/AzureTTS.ts` around lines 164 - 196, The code in
AzureTTS.ts constructs outputFormat via mapFormat(...) but still returns the
original requested options.format in the TTSResult, which can incorrectly label
audio when the requested format is unsupported; update the send-flow in the
synthesize method (where outputFormat is computed and the result object is
created) to validate the requested format: call mapFormat(options.format) and if
it returns a different actual output (or null/undefined for unsupported) either
throw a clear error rejecting unsupported formats or set result.format to the
actual produced format derived from outputFormat (e.g., map outputFormat back to
a canonical extension) so the returned TTSResult.format reflects the real
payload; reference mapFormat, outputFormat, and the result: TTSResult creation
to locate and fix the logic.
src/lib/adapters/stt/azureSTTHandler.ts-471-486 (1)

471-486: ⚠️ Potential issue | 🟠 Major

Use the negotiated input format for streamed audio frames.

Every chunk is tagged as audio/x-wav here, even if the caller is streaming mp3, ogg, or webm. That will make Azure decode the stream incorrectly unless WAV is the only supported input and enforced up front.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/azureSTTHandler.ts` around lines 471 - 486, The code
currently hardcodes Content-Type: audio/x-wav when sending audio frames and the
final endMessage, which breaks non-WAV streams; update the send logic in the
loop and the endMessage to use the negotiated input MIME type (e.g., a variable
like contentType or negotiatedAudioFormat obtained earlier in the handler)
instead of the literal 'audio/x-wav', ensuring both the headerBuffer/Message and
the endMessage use that variable (refer to audioStream, ws, requestId and the
send logic in azureSTTHandler.ts to locate and replace the hardcoded value).
src/lib/adapters/stt/azureSTTHandler.ts-369-381 (1)

369-381: ⚠️ Potential issue | 🟠 Major

isFinal is wrong when interim results are enabled.

This makes every successful segment non-final whenever interimResults is true, but the handler never emits a corresponding final segment. Downstream voice agents won't know when an utterance is actually complete.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/azureSTTHandler.ts` around lines 369 - 381, The isFinal
flag is computed incorrectly as "!azureOptions.interimResults", causing
interim-mode segments never to be marked final; change the logic in the
TranscriptionSegment construction (in azureSTTHandler) so that when
interimResults is true you set isFinal based on the recognition result (e.g.,
data.RecognitionStatus === "Success"), and when interimResults is false you mark
segments final (true). Update the isFinal assignment (referencing azureOptions,
data.RecognitionStatus, and TranscriptionSegment) accordingly so final segments
are emitted when the service reports Success.
test/continuous-test-suite-voice.ts-333-335 (1)

333-335: ⚠️ Potential issue | 🟠 Major

The stream-handler export checks are also unconditional passes.

recordTest(..., exists || true, !exists) means this block never fails when an expected export disappears. The suite should fail here, not silently downgrade the miss to a skip.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-voice.ts` around lines 333 - 335, The export
existence checks in the loop use "exists || true" which makes the test always
pass; update the call to recordTest in the loop over expectedExports so the
second argument is the actual boolean "exists" (not short-circuited) and the
skip flag is not used to silently downgrade failures (replace the third argument
so it does not mark a missing export as skipped — e.g., pass false or remove the
skip flag), locating the change in the loop that references expectedExports,
streamHandler, and recordTest.
test/continuous-test-suite-voice.ts-48-53 (1)

48-53: 🛠️ Refactor suggestion | 🟠 Major

Use a type alias for TestResult.

interface is forbidden in this repo; switch this to type TestResult = { ... }.

As per coding guidelines "Never use interface. Always use type X = { ... }."

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-voice.ts` around lines 48 - 53, Replace the
forbidden interface declaration with a type alias: change the `interface
TestResult { ... }` to `type TestResult = { name: string; passed: boolean;
skipped?: boolean; error?: string; }`, preserving all property names and
optional markers exactly and updating any imports/usages if your editor flags
them; ensure the symbol `TestResult` remains exported/visible as before.
test/continuous-test-suite-voice.ts-347-356 (1)

347-356: ⚠️ Potential issue | 🟠 Major

Import the canonical voice types module here.

This check is pointed at ../src/lib/voice/types.js, but the PR moves shared voice types into src/lib/types/voice.ts. In its current form the suite will skip the real type-export check instead of exercising the canonical module.

As per coding guidelines "src/lib/types/**/*.ts: All type definitions must go in src/lib/types/. Never create type files inside feature subdirectories."

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-voice.ts` around lines 347 - 356, The test
currently imports "../src/lib/voice/types.js" which is outdated; update the
dynamic import to the canonical types module moved to src/lib/types by importing
"../src/lib/types/voice.js" (so the runtime .js path is used), keep the existing
error handling and recordTest calls (symbols: the dynamic import expression and
recordTest) so the suite actually exercises the shared type-export module
instead of the old feature-local path.
src/cli/commands/voice.ts-359-475 (1)

359-475: 🛠️ Refactor suggestion | 🟠 Major

Move these flag definitions back into the CLI command factory.

This command group is defining its option schema inline instead of consuming the centralized CLI option definitions. That makes the new voice commands drift-prone relative to the rest of the CLI surface.

Based on learnings "All CLI command options and flag definitions are centralized in src/cli/factories/commandFactory.ts."

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/cli/commands/voice.ts` around lines 359 - 475, The voice command group is
defining option schemas inline in createVoiceCommands (affecting the
"synthesize", "transcribe", and "providers" subcommands) instead of using the
centralized CLI option definitions; refactor createVoiceCommands to import and
reuse the shared option definitions from src/cli/factories/commandFactory.ts
(e.g., the centralized voice/tts/stt option objects or helper like
getOption/optionFactory) and replace the .option/.positional inline calls for
flags like "provider", "voice", "output", "format", "speed", "pitch", "play",
"file", "language", "diarization", "word-timestamps", and "type" so
handleSynthesize, handleTranscribe, and handleProviders receive the standardized
args shape; ensure you update imports and types in createVoiceCommands
accordingly and remove the duplicated inline schemas.
src/lib/voice/providers/OpenAITTS.ts-134-176 (1)

134-176: ⚠️ Potential issue | 🟠 Major

Don't silently coerce unsupported output formats to mp3.

mapFormat() only handles four formats, but TTSResult.format still echoes options.format at Line 172. If a caller requests one of the newly added formats, the request is sent as mp3 while the SDK reports the original format, which will produce mislabeled files and confusing playback failures.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/OpenAITTS.ts` around lines 134 - 176, The code sends
the request using the mapped responseFormat but returns TTSResult.format as the
original options.format, causing mislabeled outputs; update OpenAITTS.mapFormat
usage so the resolved format is used in the result (e.g., set result.format to
the mapped responseFormat or a validated canonical format) and ensure sample
rate is derived from that canonical format via getSampleRate; alternatively
validate options.format up-front in the method (using mapFormat) and throw an
error for unsupported formats instead of silently coercing to "mp3". Ensure
references: mapFormat, responseFormat, TTSResult, options.format, and
getSampleRate are updated accordingly.
test/continuous-test-suite-voice.ts-122-131 (1)

122-131: ⚠️ Potential issue | 🟠 Major

Align the TTS provider matrix with the providers this PR actually ships.

src/lib/voice/index.ts:146-197 only exports Google, ElevenLabs, OpenAI, and Azure TTS providers, but this list also includes Polly/PlayHT/Deepgram/Cartesia and treats missing modules as SKIP. That makes the suite pass even when the expected provider set is wrong.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-voice.ts` around lines 122 - 131, The ttsProviders
matrix in test/continuous-test-suite-voice.ts includes providers not exported by
src/lib/voice/index.ts, causing tests to silently SKIP extras; update the
ttsProviders array (symbol: ttsProviders) to only include the providers actually
exported (GoogleTTS, ElevenLabsTTS, OpenAITTS, AzureTTS) and remove PollyTTS,
PlayHTTTS, DeepgramTTS, CartesiaTTS entries, ensuring the envKey values match
the exported providers' expected env vars and the test no longer masks
missing-module failures.
src/lib/voice/providers/GoogleTTS.ts-90-92 (1)

90-92: ⚠️ Potential issue | 🟠 Major

Don't report service-account auth as supported until it's implemented.

isConfigured() returns true when only credentialsPath is set, but getAccessToken() always returns an empty string after logging a warning. That sends callers down a “configured” path that can only 401.

🔧 Minimal safe fix
  isConfigured(): boolean {
-    return this.apiKey !== null || this.credentialsPath !== null;
+    return this.apiKey !== null;
  }

Also applies to: 127-129, 232-234, 371-375

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/GoogleTTS.ts` around lines 90 - 92, isConfigured()
and similar checks currently return true when credentialsPath is set even though
getAccessToken() only logs a warning and returns an empty string; change those
checks (e.g., isConfigured(), anywhere using credentialsPath like the other
occurrences referenced) to only report configured/supported when a usable auth
method exists (apiKey is non-null) or implement real service-account token
retrieval; specifically, update isConfigured() to require this.apiKey (and
update the other identical checks) so callers aren't misled into a configured
path that will 401, or alternatively implement getAccessToken() to actually
exchange the service account credentialsPath for a valid token before leaving
credentialsPath treated as supported.
src/lib/adapters/stt/deepgramSTTHandler.ts-95-105 (1)

95-105: ⚠️ Potential issue | 🟠 Major

Keep SUPPORTED_FORMATS and getContentType() in sync.

mp4, mpeg, and mpga are advertised as supported on Lines 95-105, but Lines 538-554 fall back to audio/wav for all three. Requests for those formats will be mislabeled on upload.

Also applies to: 538-554

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 95 - 105,
SUPPORTED_FORMATS lists "mp4", "mpeg", and "mpga" but getContentType() currently
falls back to "audio/wav" for those, causing incorrect Content-Type on upload;
update the getContentType() implementation to return the proper MIME types for
these extensions (e.g., "mp4" -> "audio/mp4" and both "mpeg" and "mpga" ->
"audio/mpeg") and ensure any fallback remains only for unknown extensions so
SUPPORTED_FORMATS and getContentType() stay in sync.
src/lib/voice/providers/GoogleTTS.ts-84-87 (1)

84-87: ⚠️ Potential issue | 🟠 Major

Use the TTS-specific env var here.

This constructor reads GOOGLE_API_KEY, but this provider is the Google Cloud TTS integration. A setup that only defines GOOGLE_TTS_API_KEY will look unconfigured unless callers pass the key explicitly.

🔧 Proposed fix
-    this.apiKey = apiKey ?? process.env.GOOGLE_API_KEY ?? null;
+    this.apiKey = apiKey ?? process.env.GOOGLE_TTS_API_KEY ?? null;

Based on learnings In the NeuroLink TTS SDK (src/lib/tts/), use GOOGLE_TTS_API_KEY environment variable specifically for Google Cloud Text-to-Speech access. GOOGLE_AI_API_KEY does not provide TTS access and should not be used for TTS functionality.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/GoogleTTS.ts` around lines 84 - 87, The constructor
in GoogleTTS.ts sets this.apiKey from GOOGLE_API_KEY which is incorrect for the
Google Cloud TTS provider; update the precedence in the GoogleTTS constructor so
it uses the explicit parameter first, then process.env.GOOGLE_TTS_API_KEY, and
then null (do not use GOOGLE_AI_API_KEY or GOOGLE_API_KEY for TTS), while
leaving this.credentialsPath to continue using
process.env.GOOGLE_APPLICATION_CREDENTIALS as before.
src/lib/voice/providers/ElevenLabsTTS.ts-68-130 (1)

68-130: ⚠️ Potential issue | 🟠 Major

Honor languageCode in getVoices(), or drop the parameter.

getVoices("ja") still returns the full voice list. Right now the argument only bypasses the cache, and every mapped voice gets the same hardcoded languageCodes, so filtered discovery is inaccurate.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/ElevenLabsTTS.ts` around lines 68 - 130, The
getVoices(languageCode?) implementation currently ignores the requested language
and hardcodes languageCodes and languageCode values; update getVoices to honor
the languageCode param by deriving each voice's supported languages from the
ElevenLabs response (use fields on ElevenLabsVoicesResponse.voice such as
language(s)/labels if present), set voice.languageCode dynamically (e.g.,
primary lang) and voice.languageCodes from the actual voice metadata, then
filter the returned voices to only those that include the requested
languageCode; also adjust caching so either cache per-language (keyed by
languageCode) or cache the full list and apply the language filter at return
time, and continue to use mapGender(…) and voice.voice_id/voice.name as before.
src/lib/voice/providers/OpenAISTT.ts-194-196 (1)

194-196: ⚠️ Potential issue | 🟠 Major

Normalize locale tags before sending language to OpenAI's Whisper API.

The language parameter must use ISO-639-1 format (2-letter codes like en). OpenAI's transcription API does not accept locale tags like en-US or pt-BR. Strip to the base language code before appending.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/OpenAISTT.ts` around lines 194 - 196, The code
currently appends options.language directly to the multipart form in OpenAISTT
(where formData.append("language", options.language) is called); normalize the
locale by extracting the base ISO-639-1 code (take substring before '-' or '_'
and lowercase) and validate it is a 2-letter alpha code before appending. Update
the logic in the OpenAISTT method that builds formData to transform
options.language -> baseLang = options.language.split(/[-_]/)[0].toLowerCase()
and only call formData.append("language", baseLang) if baseLang matches
/^[a-z]{2}$/.
src/lib/adapters/stt/googleSTTHandler.ts-184-189 (1)

184-189: ⚠️ Potential issue | 🟠 Major

Capability claims "streaming" but transcribeStream is not implemented.

getCapabilities() returns ["stt", "streaming"], but this class has no transcribeStream method. Consumers relying on capability checks will incorrectly assume streaming is available.

🔧 Option 1: Remove streaming capability
   getCapabilities(): VoiceCapability[] {
-    return ["stt", "streaming"];
+    return ["stt"];
   }
🔧 Option 2: Add placeholder streaming method
async *transcribeStream(
  _audioStream: AsyncIterable<Buffer>,
  _options: STTOptions,
): AsyncIterable<TranscriptionSegment> {
  throw new Error("Streaming not implemented for google-stt adapter");
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/googleSTTHandler.ts` around lines 184 - 189,
getCapabilities claims "streaming" but there is no transcribeStream implemented;
add an explicit placeholder async generator method transcribeStream in the
googleSTT handler class (signature: async *transcribeStream(_audioStream:
AsyncIterable<Buffer>, _options: STTOptions):
AsyncIterable<TranscriptionSegment>) that immediately throws a clear
Error("Streaming not implemented for google-stt adapter") so consumers see that
streaming is unavailable, or alternatively remove "streaming" from
getCapabilities() if you prefer to signal no streaming support; reference
getCapabilities and transcribeStream when making the change.
src/lib/voice/providers/GoogleSTT.ts-482-492 (1)

482-492: ⚠️ Potential issue | 🟠 Major

Placeholder getAccessToken will cause silent auth failures.

When credentialsPath is set but apiKey is not, the code path at lines 271-273 calls getAccessToken(), which returns an empty string. This will send Authorization: Bearer headers, causing 401 errors. Either implement service account auth or throw an error indicating it's unsupported.

🔧 Proposed fix to throw instead of returning empty
   private async getAccessToken(): Promise<string> {
-    // In production, this would use the Google Auth library
-    // For now, return empty (API key should be used instead)
-    logger.warn(
-      "[GoogleSTTHandler] Service account auth not implemented, use API key",
-    );
-    return "";
+    throw new Error(
+      "Service account authentication not implemented. Please use GOOGLE_API_KEY instead.",
+    );
   }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/GoogleSTT.ts` around lines 482 - 492, The current
getAccessToken() in GoogleSTT returns an empty string which causes silent 401s
when credentialsPath is provided but apiKey is not; update getAccessToken() (in
class GoogleSTT/GoogleSTTHandler) to throw a descriptive Error (e.g., "Service
account auth not implemented; provide apiKey or implement service account flow")
instead of returning "" so the caller fails fast, and ensure any callers of
getAccessToken() will surface that error (no silent Authorization: Bearer
headers).
src/lib/adapters/stt/gladiaSTTHandler.ts-316-322 (1)

316-322: ⚠️ Potential issue | 🟠 Major

WebSocket with headers won't work in Node.js without dynamic import.

The code uses the global WebSocket constructor with a headers option. In Node.js, there is no global WebSocket, and even if polyfilled, the standard WebSocket API doesn't support custom headers. Other streaming handlers in this PR (e.g., DeepgramSTT) dynamically import the ws package. This will throw a ReferenceError in Node.js.

🔧 Proposed fix using dynamic import
+    // Import WebSocket for Node.js
+    const { default: WebSocket } = await import("ws");
+
     // Create WebSocket connection
     const ws = new WebSocket(wsUrl, {
-      // `@ts-expect-error` - headers are supported by Node.js WebSocket libraries
       headers: {
         "x-gladia-key": this.apiKey,
       },
     });
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/gladiaSTTHandler.ts` around lines 316 - 322, The
WebSocket instantiation using the global WebSocket with a headers option will
fail in Node.js; update the ws creation in gladiaSTTHandler (the block that
creates const ws = new WebSocket(wsUrl, { headers: { "x-gladia-key": this.apiKey
} })) to dynamically import the 'ws' package at runtime (e.g., const {
WebSocket: NodeWebSocket } = await import('ws')) and then instantiate
NodeWebSocket with wsUrl and the headers option so Node supports custom headers;
ensure any types/ts-expect-error are removed or adjusted and that wsUrl and
this.apiKey are passed to the imported WebSocket constructor.
src/lib/voice/providers/OpenAIRealtime.ts-437-456 (1)

437-456: ⚠️ Potential issue | 🟠 Major

Missing call_id from the event breaks function-call response.

The response.function_call_arguments.done event includes a call_id field that must be passed back in the function_call_output item. Using the function name instead (line 508) will cause the OpenAI Realtime API to reject or misroute the response.

🔧 Proposed fix to capture and use `call_id`
       case "response.function_call_arguments.done": {
         const funcEvent = event as {
           name?: string;
           arguments?: string;
+          call_id?: string;
         };
-        if (funcEvent.name && funcEvent.arguments) {
+        if (funcEvent.name && funcEvent.arguments && funcEvent.call_id) {
           try {
             const args = JSON.parse(funcEvent.arguments) as Record<
               string,
               unknown
             >;
-            this.handleFunctionCall(funcEvent.name, args);
+            this.handleFunctionCall(funcEvent.name, args, funcEvent.call_id);
           } catch {

And update handleFunctionCall:

   private async handleFunctionCall(
     name: string,
     args: Record<string, unknown>,
+    callId: string,
   ): Promise<void> {
     try {
       const result = await this.emitFunctionCall(name, args);

       // Send function result back
       if (this.ws && this.isConnected()) {
         this.ws.send(
           JSON.stringify({
             type: "conversation.item.create",
             item: {
               type: "function_call_output",
-              call_id: name, // Note: This should be the actual call_id from the event
+              call_id: callId,
               output: JSON.stringify(result),
             },
           }),
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/OpenAIRealtime.ts` around lines 437 - 456, The
handler for the "response.function_call_arguments.done" event currently reads
only funcEvent.name and funcEvent.arguments and calls handleFunctionCall(name,
args); update it to also read funcEvent.call_id (or callId) from the incoming
event and pass that call_id through so the subsequent function_call_output uses
the original call_id rather than the function name; specifically, in the case
block for "response.function_call_arguments.done" extract call_id from the event
object (alongside name and arguments), parse arguments into args as before, and
call handleFunctionCall(funcEvent.name, args, funcEvent.call_id) or otherwise
forward the call_id to the code that constructs the function_call_output item so
the outgoing item includes call_id.
src/lib/neurolink.ts-11805-11813 (1)

11805-11813: ⚠️ Potential issue | 🟠 Major

These convenience APIs bypass the configured voice providers.

If the caller omits provider, these methods hard-code elevenlabs / deepgram instead of honoring the provider already configured via initializeVoice(). That makes streaming and voice discovery drift from the SDK's active voice configuration.

Also applies to: 11905-11908, 12091-12094

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 11805 - 11813, The streaming convenience
APIs (e.g., synthesizeStream) hard-code a fallback provider
("elevenlabs"/"deepgram") when caller omits provider, which bypasses the voice
provider configured via initializeVoice(); update these calls to default to the
SDK's active provider instead of a string literal by calling the VoiceFactory
method that returns the currently configured provider (e.g., use
VoiceFactory.getActiveProvider() or VoiceFactory.getConfiguredProvider() as
available) and pass that into VoiceFactory.createTTSProvider/createSTTProvider;
make the same change for the other affected streaming methods that currently use
provider ?? "elevenlabs" or provider ?? "deepgram" so they respect the
initialized voice/STT configuration.
src/lib/neurolink.ts-11746-11753 (1)

11746-11753: ⚠️ Potential issue | 🟠 Major

Forward the full CompositeVoiceConfig.

initializeVoice() accepts CompositeVoiceConfig, but this constructor call drops ttsOptions / sttOptions plus streaming and latencyMode. Those caller-supplied settings are silently ignored.

🧩 Proposed fix
     this.compositeVoice = new CompositeVoice({
       ttsProvider: config.ttsProvider,
       sttProvider: config.sttProvider,
       defaultTTSOptions: config.defaultTTSOptions,
       defaultSTTOptions: config.defaultSTTOptions,
+      ttsOptions: config.ttsOptions,
+      sttOptions: config.sttOptions,
       trackHistory: config.trackHistory ?? true,
       maxHistoryTurns: config.maxHistoryTurns ?? 50,
+      streaming: config.streaming,
+      latencyMode: config.latencyMode,
     });
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 11746 - 11753, The CompositeVoice is being
constructed with only a subset of fields, dropping caller-supplied ttsOptions,
sttOptions, streaming and latencyMode from the CompositeVoiceConfig; update the
constructor call that creates new CompositeVoice(...) so it forwards the entire
CompositeVoiceConfig (e.g., pass config or explicitly include ttsOptions,
sttOptions, streaming, latencyMode along with ttsProvider, sttProvider,
defaultTTSOptions, defaultSTTOptions, trackHistory and maxHistoryTurns) to
preserve caller settings and keep any defaults/nullable handling the same in
initializeVoice/CompositeVoice.
src/lib/neurolink.ts-11931-11950 (1)

11931-11950: ⚠️ Potential issue | 🟠 Major

Fallback transcription can yield an empty stream.

When the provider lacks transcribeStream(), this branch only re-emits result.segments. Providers that return plain result.text without segment metadata will produce no items, so the transcription is lost.

🎙️ Proposed fix
-      if (result.segments) {
+      if (result.segments?.length) {
         for (const segment of result.segments) {
           yield {
             ...segment,
             start: segment.start ?? segment.startTime ?? 0,
             end: segment.end ?? segment.endTime ?? 0,
             confidence: segment.confidence ?? 0,
           } as TranscriptionSegment;
         }
+      } else if (result.text) {
+        yield {
+          text: result.text,
+          start: 0,
+          end: 0,
+          confidence: result.confidence ?? 0,
+          isFinal: true,
+        } as TranscriptionSegment;
       }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 11931 - 11950, The fallback branch that
collects audio and calls sttProvider.transcribe only yields result.segments,
which drops transcriptions when the provider returns only result.text; update
the branch in neurolink.ts (the block using audioStream, sttProvider.transcribe,
result.segments and TranscriptionSegment) to handle non-segment results by
yielding a single TranscriptionSegment when segments are absent: construct a
segment with text from result.text (or result.transcript), start = 0, end = 0
(or best-effort duration if available), and confidence = result.confidence ?? 0,
preserving existing behavior for providers that do return result.segments.
src/lib/neurolink.ts-11741-11745 (1)

11741-11745: ⚠️ Potential issue | 🟠 Major

Don't log raw voice config.

CompositeVoiceConfig can carry provider config objects, so logging config here can leak credentials into debug logs. Guard the serialization and sanitize it first.

🔒 Proposed fix
-    logger.debug("[NeuroLink] Initializing voice capabilities", { config });
+    if (logger.shouldLog("debug")) {
+      logger.debug("[NeuroLink] Initializing voice capabilities", {
+        config: transformParamsForLogging(
+          config as unknown as Record<string, unknown>,
+        ),
+      });
+    }
As per coding guidelines: use `logger.shouldLog("debug")` to guard expensive serialization before logging and `transformParamsForLogging()` to safely strip secrets before logging.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 11741 - 11745, The debug log in
initializeVoice currently logs the raw CompositeVoiceConfig which may contain
secrets; change it to first check logger.shouldLog("debug") and only then call
transformParamsForLogging(config) and log the sanitized result (e.g.,
logger.debug("[NeuroLink] Initializing voice capabilities", { config:
transformParamsForLogging(config) })), ensuring you reference the
initializeVoice method, CompositeVoiceConfig, logger.shouldLog("debug") and
transformParamsForLogging() when updating the logging call.
src/lib/neurolink.ts-12497-12567 (1)

12497-12567: ⚠️ Potential issue | 🟠 Major

validateVoiceProvider() rejects supported STT providers.

This classifier never routes google-stt or azure-stt through VoiceFactory.createSTTProvider(), so those providers currently fall through to Unknown provider.

🛠️ Proposed fix
-      } else if (
-        ["deepgram", "gladia", "whisper", "assemblyai"].includes(provider)
-      ) {
+      } else if (
+        [
+          "deepgram",
+          "gladia",
+          "whisper",
+          "assemblyai",
+          "google-stt",
+          "azure-stt",
+        ].includes(provider)
+      ) {
         const sttProvider = await VoiceFactory.createSTTProvider(
-          provider as "deepgram" | "gladia" | "whisper",
+          provider as
+            | "deepgram"
+            | "gladia"
+            | "whisper"
+            | "assemblyai"
+            | "google-stt"
+            | "azure-stt",
         );
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 12497 - 12567, validateVoiceProvider
currently treats "google-stt" and "azure-stt" as unknown because the STT branch
only checks ["deepgram","gladia","whisper","assemblyai"]; update that branch to
include "google-stt" and "azure-stt" in the array and in the TypeScript cast for
VoiceFactory.createSTTProvider (e.g., add "google-stt" | "azure-stt" to the
union), so createSTTProvider(...) will be called for those providers and their
validateConfig path is executed.
src/lib/voice/providers/GeminiLive.ts-446-476 (1)

446-476: ⚠️ Potential issue | 🟠 Major

Function call errors are logged but not propagated to caller.

In handleFunctionCall(), errors are caught and logged (Lines 471-474) but the function result is not sent back to Gemini on failure. The model may hang waiting for a response.

Suggested fix - send error response to Gemini
     } catch (err: unknown) {
       logger.error(
         `[GeminiLiveHandler] Function call failed: ${err instanceof Error ? err.message : String(err)}`,
       );
+      // Send error response to Gemini so it doesn't hang
+      if (this.ws && this.isConnected()) {
+        const errorResponse = {
+          toolResponse: {
+            functionResponses: [
+              {
+                id: callId,
+                name,
+                response: { error: err instanceof Error ? err.message : String(err) },
+              },
+            ],
+          },
+        };
+        this.ws.send(JSON.stringify(errorResponse));
+        this.pendingFunctionCalls.delete(callId);
+      }
     }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/GeminiLive.ts` around lines 446 - 476,
handleFunctionCall currently logs errors but never replies to Gemini, which can
leave the model waiting; in the catch block for handleFunctionCall, send a
toolResponse over this.ws (after verifying this.ws && this.isConnected()) that
mirrors the successful response shape but contains an error payload (e.g., { id:
callId, name, error: { message: errMessage, type: errType? } }), JSON.stringify
it, and call this.pendingFunctionCalls.delete(callId) so the pending call is
cleared; retain the existing logger.error call and ensure you guard the send
behind the same connection check used elsewhere (this.ws && this.isConnected())
and include the callId and name in the error response so Gemini can correlate
it.
🟡 Minor comments (8)
test/continuous-test-suite-voice.ts-84-90 (1)

84-90: ⚠️ Potential issue | 🟡 Minor

This realtime-provider assertion never fails.

realtimeProviders.length >= 0 is always true, so this check won't catch a broken registry. Use > 0 if the intent is to verify at least one realtime provider is wired.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-voice.ts` around lines 84 - 90, The test asserts
realtimeProviders.length >= 0 which is always true; change the assertion in the
test using factory1.getProvidersByType("realtime") and recordTest to require at
least one provider by using realtimeProviders.length > 0 (i.e., update the call
that currently does recordTest("Realtime providers available",
realtimeProviders.length >= 0) to use > 0 so the check will fail when no
realtime providers are registered).
src/lib/adapters/stt/deepgramSTTHandler.ts-318-323 (1)

318-323: ⚠️ Potential issue | 🟡 Minor

This appears to be dead code—the active implementation is in src/lib/voice/providers/DeepgramSTT.ts which correctly uses the ws npm package.

The technical concern is valid: the built-in WHATWG WebSocket constructor does not support a { headers: ... } options object as shown (the second argument is reserved for subprotocols only). However, this file (deepgramSTTHandler.ts) is not imported or used anywhere in the codebase. The actual streaming implementation in DeepgramSTT.ts uses import("ws") which does support headers, so authentication works correctly there.

If this file is intended to be kept as an alternative implementation, update it to use the ws package or an alternative auth method (query parameters, cookies, or post-connection auth). Otherwise, remove it to avoid confusion.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 318 - 323, This file
contains dead/incorrect code using the WHATWG WebSocket constructor with a
headers option (const ws = new WebSocket(wsUrl, { headers: { Authorization:
`Token ${this.apiKey}` } })), which doesn’t support headers; either remove the
unused deepgramSTTHandler.ts entirely to avoid confusion, or update the
implementation to use the ws npm package like in DeepgramSTT.ts (import or
dynamic import of "ws") and construct the socket using new Ws(wsUrl, { headers:
{ Authorization: `Token ${this.apiKey}` } }) and adjust types accordingly so
authentication works as intended; ensure you update any exports/usage so the
codebase references the correct implementation.
src/lib/voice/providers/GoogleSTT.ts-88-91 (1)

88-91: ⚠️ Potential issue | 🟡 Minor

supportsStreaming = true is misleading for a chunked-batch implementation.

The transcribeStream method buffers ~5 seconds of audio and calls the synchronous transcribe API repeatedly. This isn't true streaming (it has 5s latency per segment and no interim results). Consider setting supportsStreaming = false or documenting this limitation clearly.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/GoogleSTT.ts` around lines 88 - 91, The
supportsStreaming flag on GoogleSTT is misleading because transcribeStream
buffers ~5s and calls the synchronous transcribe method (no true low-latency
streaming or interim results); update the implementation by either setting the
public readonly supportsStreaming property to false on GoogleSTT or explicitly
document the limitation in the class/method comments and README, and add a note
in the transcribeStream method header explaining it performs chunked-batch
buffering (~5s latency) and does not provide interim results; reference the
supportsStreaming property and transcribeStream method when making the change.
src/lib/voice/providers/DeepgramSTT.ts-526-530 (1)

526-530: ⚠️ Potential issue | 🟡 Minor

Missing connection timeout for WebSocket.

The await new Promise for connection open has no timeout. If the WebSocket fails to connect (network issues, wrong URL), the method hangs indefinitely. Other handlers in this PR use timeouts (e.g., GladiaSTTHandler uses 10s).

🔧 Proposed fix with timeout
     // Wait for connection
     await new Promise<void>((resolve, reject) => {
+      const timeout = setTimeout(() => {
+        ws.close();
+        reject(STTError.streamError("WebSocket connection timeout", "deepgram"));
+      }, 10000);
+
-      ws.on("open", () => resolve());
-      ws.on("error", reject);
+      ws.on("open", () => {
+        clearTimeout(timeout);
+        resolve();
+      });
+      ws.on("error", (err) => {
+        clearTimeout(timeout);
+        reject(err);
+      });
     });
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/DeepgramSTT.ts` around lines 526 - 530, The
connection wait for the WebSocket (ws) currently awaits forever; add a timeout
like other handlers (e.g., GladiaSTTHandler's 10s) so the Promise rejects if
"open" doesn't occur in time. Modify the Promise that listens to ws.on("open")
and ws.on("error") to also set a timer (clear it on open/error) which rejects
after 10_000ms, and ensure any created timer is cleaned up to avoid leaks;
update the calling method in DeepgramSTT (the block that awaits new Promise for
ws) to handle the rejection accordingly.
src/lib/voice/providers/OpenAIRealtime.ts-448-448 (1)

448-448: ⚠️ Potential issue | 🟡 Minor

Async function call handler invoked without await.

handleFunctionCall is async but is called synchronously in the switch case. Any errors will become unhandled promise rejections instead of being caught and emitted via emitError.

🛠️ Proposed fix
-            this.handleFunctionCall(funcEvent.name, args);
+            void this.handleFunctionCall(funcEvent.name, args).catch((err) => {
+              this.emitError(err instanceof Error ? err : new Error(String(err)));
+            });
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/OpenAIRealtime.ts` at line 448, The switch case is
invoking the async method handleFunctionCall(funcEvent.name, args) without
awaiting it, causing unhandled promise rejections; modify the caller (the switch
handler) to either await this.handleFunctionCall(...) or explicitly attach a
catch that forwards errors to this.emitError(err) — ensure the enclosing
function is marked async if you add await, or use
this.handleFunctionCall(...).catch(err => this.emitError(err)) so all errors are
caught and emitted.
src/lib/types/voice.ts-900-912 (1)

900-912: ⚠️ Potential issue | 🟡 Minor

Type guard incorrectly requires optional property.

isTranscriptionSegment checks that obj.index === "number" (Line 908), but TranscriptionSegment.index is optional (index?: number at Line 201). Valid segments without an index will fail this guard.

Suggested fix
 export function isTranscriptionSegment(
   value: unknown,
 ): value is TranscriptionSegment {
   if (!value || typeof value !== "object") {
     return false;
   }
   const obj = value as Record<string, unknown>;
   return (
-    typeof obj.index === "number" &&
+    (obj.index === undefined || typeof obj.index === "number") &&
     typeof obj.text === "string" &&
     typeof obj.isFinal === "boolean"
   );
 }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/types/voice.ts` around lines 900 - 912, isTranscriptionSegment
incorrectly requires the optional property index to be present; update the guard
in isTranscriptionSegment so it accepts objects where obj.index is either
undefined or a number (e.g., replace the strict typeof obj.index === "number"
check with a conditional that allows undefined or a number), keeping the
existing checks for obj.text and obj.isFinal so the function still narrows to
TranscriptionSegment when appropriate.
src/lib/voice/voiceAgent.ts-68-69 (1)

68-69: ⚠️ Potential issue | 🟡 Minor

Readonly property is mutated.

this.config is declared readonly at Line 68, but updateVoiceSettings() mutates this.config.voiceSettings at Lines 469-472. Either remove the readonly modifier or use a separate mutable settings field.

Proposed fix

Option 1 - Remove readonly (if mutation is intended):

-  private readonly config: VoiceAgentConfig;
+  private config: VoiceAgentConfig;

Option 2 - Use a separate mutable field:

   private readonly config: VoiceAgentConfig;
+  private voiceSettingsOverrides: Partial<VoiceAgentConfig["voiceSettings"]> = {};

Also applies to: 465-475

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/voiceAgent.ts` around lines 68 - 69, The property this.config
(type VoiceAgentConfig) is declared readonly but updateVoiceSettings mutates
this.config.voiceSettings; fix by either removing the readonly modifier from the
config declaration so mutations are allowed, or keep config readonly and
introduce a new mutable field (e.g., mutableVoiceSettings or voiceSettings) that
updateVoiceSettings reads/writes instead; update all references in
updateVoiceSettings and any other methods that modify or rely on voiceSettings
to use the chosen mutable field, leaving the original this.config as an
immutable configuration object.
src/lib/voice/voiceFactory.ts-554-576 (1)

554-576: ⚠️ Potential issue | 🟡 Minor

Static has*Provider methods may return false before initialization completes.

The hasTTSProvider, hasSTTProvider, and hasRealtimeProvider methods call getType() which reads from typeMap synchronously. If called before initialization completes, they will incorrectly return false.

Suggested fix - add async variants or document limitation
   /**
    * Check if a TTS provider exists (static convenience method)
+   * NOTE: Returns false if factory is not yet initialized.
+   * Use ensureInitialized() first for reliable results.
    */
   static hasTTSProvider(nameOrAlias: string): boolean {

Or add async variants:

static async hasTTSProviderAsync(nameOrAlias: string): Promise<boolean> {
  const factory = VoiceFactory.getInstance();
  await factory.ensureInitialized();
  return factory.getType(nameOrAlias) === "tts";
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/voiceFactory.ts` around lines 554 - 576, The static
hasTTSProvider/hasSTTProvider/hasRealtimeProvider methods call getType()
synchronously and can return false if initialization hasn't completed; to fix,
add async variants (e.g., hasTTSProviderAsync, hasSTTProviderAsync,
hasRealtimeProviderAsync) that call
VoiceFactory.getInstance().ensureInitialized() before using getType(), or
alternatively make the existing static methods await ensureInitialized() and
return Promise<boolean>; reference the existing methods
hasTTSProvider/hasSTTProvider/hasRealtimeProvider, the instance method getType,
and the initialization helper ensureInitialized when implementing the change so
callers get correct results after initialization.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6cb9af04-3973-4fe4-82dc-719ede7a4c8c

📥 Commits

Reviewing files that changed from the base of the PR and between 49f56cd and a118f5e.

📒 Files selected for processing (33)
  • src/cli/commands/voice.ts
  • src/cli/parser.ts
  • src/lib/adapters/stt/assemblyaiSTTHandler.ts
  • src/lib/adapters/stt/azureSTTHandler.ts
  • src/lib/adapters/stt/deepgramSTTHandler.ts
  • src/lib/adapters/stt/gladiaSTTHandler.ts
  • src/lib/adapters/stt/googleSTTHandler.ts
  • src/lib/adapters/stt/whisperSTTHandler.ts
  • src/lib/neurolink.ts
  • src/lib/types/index.ts
  • src/lib/types/tts.ts
  • src/lib/types/voice.ts
  • src/lib/voice/RealtimeVoiceAPI.ts
  • src/lib/voice/STTProvider.ts
  • src/lib/voice/audio-utils.ts
  • src/lib/voice/compositeVoice.ts
  • src/lib/voice/errors.ts
  • src/lib/voice/index.ts
  • src/lib/voice/providers/AzureSTT.ts
  • src/lib/voice/providers/AzureTTS.ts
  • src/lib/voice/providers/DeepgramSTT.ts
  • src/lib/voice/providers/ElevenLabsTTS.ts
  • src/lib/voice/providers/GeminiLive.ts
  • src/lib/voice/providers/GoogleSTT.ts
  • src/lib/voice/providers/GoogleTTS.ts
  • src/lib/voice/providers/OpenAIRealtime.ts
  • src/lib/voice/providers/OpenAISTT.ts
  • src/lib/voice/providers/OpenAITTS.ts
  • src/lib/voice/stream-handler.ts
  • src/lib/voice/voiceAgent.ts
  • src/lib/voice/voiceFactory.ts
  • src/lib/voice/voiceRegistry.ts
  • test/continuous-test-suite-voice.ts

Comment thread src/lib/adapters/stt/assemblyaiSTTHandler.ts Outdated
Comment thread src/lib/voice/providers/AzureSTT.ts Outdated
@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@murdore
murdore force-pushed the feat/voice-speech-integration branch from 5cc6346 to f4cd22b Compare April 18, 2026 19:42
@murdore

murdore commented Apr 18, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@murdore

murdore commented Apr 18, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Full review triggered.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 9

Note

Due to the large number of review comments, Critical severity comments were prioritized as inline comments.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
src/cli/factories/commandFactory.ts (1)

318-323: ⚠️ Potential issue | 🟡 Minor

Keep CLI TTS formats in sync with AudioFormat.

The public type now accepts m4a, flac, webm, mp4, mpeg, and mpga, but the CLI still rejects them and casts ttsFormat to the old four-format union.

🐛 Proposed fix
+  AudioFormat,
   AnthropicAuthConfig,
   AnthropicAuthMethod,
   ClaudeSubscriptionTier,
 } from "../../lib/types/index.js";
     ttsFormat: {
       type: "string" as const,
-      choices: ["mp3", "wav", "ogg", "opus"],
+      choices: ["mp3", "wav", "ogg", "opus", "m4a", "flac", "webm", "mp4", "mpeg", "mpga"],
       default: "mp3",
       description: "Audio output format",
     },
-      ttsFormat: argv.ttsFormat as "mp3" | "wav" | "ogg" | "opus" | undefined,
+      ttsFormat: argv.ttsFormat as AudioFormat | undefined,
               format:
-                (enhancedOptions.ttsFormat as "mp3" | "wav" | "ogg" | "opus") ||
-                undefined,
+                (enhancedOptions.ttsFormat as AudioFormat | undefined) ||
+                undefined,
             format:
-              (enhancedOptions.ttsFormat as "mp3" | "wav" | "ogg" | "opus") ||
-              undefined,
+              (enhancedOptions.ttsFormat as AudioFormat | undefined) ||
+              undefined,

Also applies to: 727-729, 2696-2698, 2966-2968

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/cli/factories/commandFactory.ts` around lines 318 - 323, The CLI option
ttsFormat is restricted to the old four values and must be updated to match the
public AudioFormat union; locate the ttsFormat option in commandFactory.ts (and
the other ttsFormat occurrences) and replace the hard-coded choices with the
full set of AudioFormat values (include
"mp3","wav","ogg","opus","m4a","flac","webm","mp4","mpeg","mpga") or, better,
derive choices from the shared AudioFormat type/enum so they stay in sync; keep
the existing default (e.g., "mp3") and ensure any casting or type annotation
accepts the expanded union.
src/lib/types/tts.ts (1)

153-159: ⚠️ Potential issue | 🟠 Major

Update runtime audio format validation for the expanded AudioFormat union.

AudioFormat now allows m4a, flac, webm, mp4, mpeg, and mpga, but VALID_AUDIO_FORMATS still rejects them via isValidTTSOptions() and isTTSResult().

🐛 Proposed fix
 export const VALID_AUDIO_FORMATS: readonly AudioFormat[] = [
   "mp3",
   "wav",
   "ogg",
   "opus",
+  "m4a",
+  "flac",
+  "webm",
+  "mp4",
+  "mpeg",
+  "mpga",
 ];
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/types/tts.ts` around lines 153 - 159, VALID_AUDIO_FORMATS is out of
sync with the expanded AudioFormat union, causing isValidTTSOptions and
isTTSResult to reject valid formats; update VALID_AUDIO_FORMATS to include the
new formats (m4a, flac, webm, mp4, mpeg, mpga) so runtime validation matches the
AudioFormat type, and ensure any logic in isValidTTSOptions() and isTTSResult()
references VALID_AUDIO_FORMATS rather than hardcoded lists so future enum
additions stay consistent.
♻️ Duplicate comments (1)
src/lib/voice/providers/AzureSTT.ts (1)

38-41: ⚠️ Potential issue | 🟠 Major

supportsStreaming still overstates Azure STT behavior.

transcribeStream() buffers ~5 seconds and repeatedly calls the batch REST endpoint, so callers do not get true continuous streaming/interim-result semantics. Either set this to false or implement Azure Speech SDK/WebSocket streaming.

Does Azure Speech REST endpoint /speech/recognition/conversation/cognitiveservices/v1 provide true continuous streaming interim transcription results, or only final REST recognition responses?

Also applies to: 282-336

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/AzureSTT.ts` around lines 38 - 41, The
supportsStreaming flag on AzureSTT incorrectly claims true streaming; update the
AzureSTT implementation by setting the public readonly supportsStreaming
property to false (or alternatively implement true streaming via the Azure
Speech SDK/WebSocket) and ensure transcribeStream()'s behavior and documentation
reflect batched REST calls (buffering ~5s and returning final results) rather
than interim streaming; reference the AzureSTT supportsStreaming property and
the transcribeStream() method and adjust any callers/tests that rely on
streaming semantics accordingly.
🟡 Minor comments (10)
src/lib/adapters/stt/deepgramSTTHandler.ts-250-256 (1)

250-256: ⚠️ Potential issue | 🟡 Minor

Hardcoded encoding/sample_rate ignores caller-provided audio format.

transcribeStream force-sets encoding=linear16 and sample_rate=16000 regardless of what the caller is actually sending through audioStream. If the audio source is 48 kHz Opus (browser MediaRecorder), 44.1 kHz PCM, or anything else, Deepgram will mis-transcribe or drop characters without any error surfacing. Read these from options (mirroring the batch path's getContentType) and fall back to sensible defaults only when unspecified.

-    queryParams.set("encoding", "linear16");
-    queryParams.set("sample_rate", "16000");
+    queryParams.set(
+      "encoding",
+      (deepgramOptions as { encoding?: string }).encoding ?? "linear16",
+    );
+    queryParams.set(
+      "sample_rate",
+      String((deepgramOptions as { sampleRate?: number }).sampleRate ?? 16000),
+    );
     queryParams.set("interim_results", "true");
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 250 - 256,
transcribeStream currently overrides caller audio format by hardcoding encoding
and sample_rate; modify transcribeStream to read encoding and sample_rate from
the provided DeepgramSTTOptions (deepgramOptions) or derive them the same way as
the batch path (use getContentType or the options fields) before calling
buildQueryParams, and only set defaults (e.g., linear16/16000) when those option
values are absent; update the code around buildQueryParams/deepgramOptions to
conditionally set queryParams.set("encoding", ...) and
queryParams.set("sample_rate", ...) from the resolved values so the WS URL
reflects the actual audio format being sent.
src/lib/adapters/stt/deepgramSTTHandler.ts-338-386 (1)

338-386: ⚠️ Potential issue | 🟡 Minor

Replace STREAMING_NOT_SUPPORTED with STREAM_ERROR for connection timeouts and mid-stream errors.

The code at lines 339 and 380 uses STT_ERROR_CODES.STREAMING_NOT_SUPPORTED for network/timeout failures, but this code semantically means "the provider doesn't support streaming." Callers branching on error codes (e.g., falling back to batch mode) will be misled into permanently disabling streaming for transient failures. Use STREAM_ERROR instead—it's the appropriate code for streaming connection and runtime failures.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 338 - 386, Replace
the incorrect error code usage: change STT_ERROR_CODES.STREAMING_NOT_SUPPORTED
to STT_ERROR_CODES.STREAM_ERROR for the STTError thrown in the WebSocket timeout
handler (the timeout callback that constructs new STTError) and for the STTError
thrown when state.errorMessage is set in the receive loop; keep the rest of the
STTError fields (category, severity, retriable, provider "deepgram") unchanged
so callers see a transient STREAM_ERROR instead of a permanent
STREAMING_NOT_SUPPORTED.
src/lib/types/voice.ts-440-499 (1)

440-499: ⚠️ Potential issue | 🟡 Minor

Add metadata for all valid AudioFormat values.

AudioFormat includes mp4, mpeg, and mpga, but AUDIO_FORMAT_DETAILS has no entries for them. Any lookup by a valid format will get undefined.

🐛 Proposed fix
   webm: {
     format: "webm",
     mimeType: "audio/webm",
     extension: ".webm",
     supportsStreaming: true,
     sampleRates: [44100, 48000],
     bitDepths: [16],
   },
+  mp4: {
+    format: "mp4",
+    mimeType: "audio/mp4",
+    extension: ".mp4",
+    supportsStreaming: false,
+    sampleRates: [44100, 48000],
+    bitDepths: [16],
+  },
+  mpeg: {
+    format: "mpeg",
+    mimeType: "audio/mpeg",
+    extension: ".mpeg",
+    supportsStreaming: true,
+    sampleRates: [8000, 16000, 22050, 24000, 44100, 48000],
+    bitDepths: [16],
+  },
+  mpga: {
+    format: "mpga",
+    mimeType: "audio/mpeg",
+    extension: ".mpga",
+    supportsStreaming: true,
+    sampleRates: [8000, 16000, 22050, 24000, 44100, 48000],
+    bitDepths: [16],
+  },
 };
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/types/voice.ts` around lines 440 - 499, AUDIO_FORMAT_DETAILS is
missing entries for valid AudioFormat values (mp4, mpeg, mpga), causing lookups
to return undefined; update the AUDIO_FORMAT_DETAILS object to include metadata
objects for "mp4", "mpeg", and "mpga" (matching the shape of AudioFormatDetails:
format, mimeType, extension, supportsStreaming, sampleRates, bitDepths) so every
AudioFormat key is present — e.g., add mp4 (format: "mp4", mimeType:
"audio/mp4", extension: ".mp4", supportsStreaming: false, sampleRates: [...],
bitDepths: [...]), mpeg (format: "mpeg", mimeType: "audio/mpeg", extension:
".mpeg" or ".mpg", supportsStreaming: true, sampleRates: [...], bitDepths:
[...]), and mpga (format: "mpga", mimeType: "audio/mpeg", extension: ".mpga",
supportsStreaming: true, sampleRates: [...], bitDepths: [...]) ensuring
sampleRates/bitDepths choices align with similar entries (e.g., mp3/opus) and
types remain Partial<Record<AudioFormat, AudioFormatDetails>>.
src/lib/neurolink.ts-11888-11968 (1)

11888-11968: ⚠️ Potential issue | 🟡 Minor

Keep createEvaluationPipeline attached to its JSDoc.

The new voice methods now sit between the evaluation pipeline docblock and createEvaluationPipeline, so generated API docs may orphan that documentation. Move the voice methods above the “Evaluation & Scoring API” section or move the createEvaluationPipeline JSDoc down to line 11968.

📝 Proposed structure
+  // ========================================
+  // Voice API
+  // ========================================
+
   /**
    * Synthesize text to speech.
    */
   async synthesize(...) { ... }

   async transcribe(...) { ... }

   async startRealtimeVoice(...) { ... }

   /**
    * Create an evaluation pipeline with the specified configuration or preset.
    * Pipelines orchestrate multiple scorers to evaluate AI responses comprehensively.
    */
   async createEvaluationPipeline(...)
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 11888 - 11968, The JSDoc for
createEvaluationPipeline has been separated from its function by the new voice
methods; to fix, ensure the createEvaluationPipeline JSDoc stays immediately
above the createEvaluationPipeline declaration by moving either the voice
methods (TTS synthesize, STT transcribe, startRealtimeVoice /
RealtimeProcessor.connect) above the "Evaluation & Scoring API" section or by
relocating the createEvaluationPipeline JSDoc block down so it directly precedes
the createEvaluationPipeline function; update references to startRealtimeVoice,
synthesize, transcribe as needed to preserve logical grouping and documentation
generation.
src/lib/voice/providers/AzureTTS.ts-226-228 (1)

226-228: ⚠️ Potential issue | 🟡 Minor

Escape SSML attribute values, not just text nodes.

voice is inserted into SSML attributes without XML attribute escaping. A malformed or user-provided voice value can break the generated SSML.

🛡️ Proposed direction
-      return azureOptions.ssmlTemplate
-        .replace("{text}", this.escapeXml(text))
-        .replace("{voice}", voice);
+      return azureOptions.ssmlTemplate
+        .replace("{text}", this.escapeXml(text))
+        .replace("{voice}", this.escapeXml(voice));
...
-  <voice name="${voice}">
+  <voice name="${this.escapeXml(voice)}">

Also applies to: 245-248

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/AzureTTS.ts` around lines 226 - 228, The SSML
assembly inserts the voice identifier into attributes without XML-attribute
escaping, which can break SSML if voice contains quotes or special chars; update
the code that builds SSML (the expression using
azureOptions.ssmlTemplate.replace("{text}",
this.escapeXml(text)).replace("{voice}", voice) and the analogous replacements
around 245-248) to escape attribute values before insertion (e.g., call a new or
existing escapeXmlAttribute/escapeXmlForAttribute helper on voice and any other
attribute substitutions) so only attribute-safe strings are injected into the
template.
src/lib/adapters/stt/assemblyaiSTTHandler.ts-405-405 (1)

405-405: ⚠️ Potential issue | 🟡 Minor

Remove the forbidden non-null assertions flagged by static analysis.

These are already under guard conditions; assign guarded locals instead of using ! so the code passes the quality gate.

♻️ Example cleanup
-        yield segments.shift()!;
+        const nextSegment = segments.shift();
+        if (nextSegment) {
+          yield nextSegment;
+        }
-            Authorization: this.apiKey!,
+            Authorization: this.apiKey,

For private methods, pass apiKey as an argument or add a local guard before building headers.

Also applies to: 432-432, 571-571, 604-604, 686-687

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/assemblyaiSTTHandler.ts` at line 405, Replace all
non-null assertion usages (e.g., the `segments.shift()!` call and other `!`
usages at the referenced locations) by assigning the potentially-null value to a
guarded local variable and performing an explicit null check before usage; for
example, `const segment = segments.shift(); if (!segment) { /* handle or
continue */ } yield segment;`. Do the same for other occurrences (lines noted
around 432, 571, 604, 686–687): assign guarded locals and
early-return/throw/continue as appropriate instead of using `!`. For private
methods that build headers and currently rely on a non-null `apiKey`, either
accept `apiKey` as a parameter or add a local guard (e.g., `const key =
this.apiKey; if (!key) throw new Error(...)`) before constructing the headers,
and use the guarded `key` variable. Ensure all replacements reference the
original symbols (`segments`, `segment`, `apiKey`, header-building helpers) so
static analysis no longer sees non-null assertions.
src/lib/voice/providers/ElevenLabsTTS.ts-82-110 (1)

82-110: ⚠️ Potential issue | 🟡 Minor

Apply the requested language filter in getVoices().

When languageCode is provided, this still returns every voice while marking languageCode: "en" for each. Consumers asking for "fr" or "de" will receive unfiltered English-labeled results.

🐛 Proposed fix
-      const voices: TTSVoice[] = data.voices.map((voice) => ({
+      let voices: TTSVoice[] = data.voices.map((voice) => ({
...
-      // Cache voices
+      if (languageCode) {
+        const requested = languageCode.toLowerCase();
+        voices = voices.filter((voice) =>
+          voice.languageCodes?.some((code) =>
+            code.toLowerCase().startsWith(requested),
+          ),
+        );
+      }
+
+      // Cache voices
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/ElevenLabsTTS.ts` around lines 82 - 110, getVoices()
currently ignores the languageCode parameter and hardcodes languageCode: "en"
and a static languageCodes list; fix by filtering data.voices to only include
voices that support the requested languageCode (e.g., check
voice.language_codes, voice.languages, or voice.labels for supported languages)
and set each TTSVoice.languageCode to the requested languageCode (or to the
voice's primary supported language if languageCode is undefined). Update the
voices mapping (referencing voice.voice_id, name, mapGender) to derive
languageCodes from the voice metadata instead of the hardcoded array, apply the
filter before creating the voices array, and keep caching behavior (voicesCache)
unchanged so caching only happens when no languageCode is passed.
src/lib/voice/RealtimeVoiceAPI.ts-155-183 (1)

155-183: ⚠️ Potential issue | 🟡 Minor

Detach realtime event handlers when connect fails.

handler.on(handlers) runs before handler.connect(). If connection fails, those callbacks remain attached and may fire during a later session.

♻️ Proposed fix
     } catch (err: unknown) {
+      if (handlers) {
+        handler.off();
+      }
+
       if (err instanceof RealtimeError) {
         throw err;
       }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/RealtimeVoiceAPI.ts` around lines 155 - 183, The handlers
currently registered via handler.on(handlers) remain attached if
handler.connect(mergedConfig) throws; update the connect flow so you unregister
those handlers on failure: after calling handler.on(handlers) and before
awaiting handler.connect, ensure the catch block calls a corresponding detach
(e.g., handler.off(handlers) or handler.removeListener/removeAllListeners as
appropriate) to remove the previously attached callbacks, then rethrow the
RealtimeError (preserving the existing RealtimeError.connectionFailed
construction); reference the existing handler.on, handler.connect and handlers
symbols and add the detach call in the catch path for non-RealtimeError
failures.
src/lib/voice/index.ts-107-122 (1)

107-122: ⚠️ Potential issue | 🟡 Minor

Export the STT providers added outside voice/providers.

AssemblyAISTTHandler is included in this PR under src/lib/adapters/stt/assemblyaiSTTHandler.ts, but it is not exposed from the consolidated voice entrypoint like the other STT providers.

♻️ Proposed export
+export { AssemblyAISTTHandler } from "../adapters/stt/assemblyaiSTTHandler.js";
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/index.ts` around lines 107 - 122, The AssemblyAISTT provider
added in src/lib/adapters/stt/assemblyaiSTTHandler.ts is not exported from the
consolidated voice entrypoint; add an export in src/lib/voice/index.ts following
the existing pattern (export the class and an alias) so the symbol AssemblyAISTT
and AssemblyAISTTHandler (or the actual exported class name from
assemblyaiSTTHandler.ts) are exposed alongside AzureSTT, DeepgramSTT, GoogleSTT,
and OpenAISTT; ensure the export path points to the
adapters/stt/assemblyaiSTTHandler module and matches the class name exported
there.
src/lib/types/realtime.ts-60-68 (1)

60-68: ⚠️ Potential issue | 🟡 Minor

JSDoc comments are misaligned with their fields.

The JSDoc lines /** Turn detection mode */ and /** Instructions/system prompt for the session */ both land above temperature?, and temperature?'s own /** Temperature for AI responses */ sits above instructions?. Tools like TSDoc / IDE tooltips will attribute each doc block to the wrong field.

🔧 Reorder to match fields
-  /** VAD threshold (0-1) */
-  vadThreshold?: number;
-  /** Turn detection mode */
-  /** Instructions/system prompt for the session */
-  /** Temperature for AI responses */
-  temperature?: number;
-  instructions?: string;
-  turnDetection?: "server_vad" | "manual";
+  /** VAD threshold (0-1) */
+  vadThreshold?: number;
+  /** Temperature for AI responses */
+  temperature?: number;
+  /** Instructions/system prompt for the session */
+  instructions?: string;
+  /** Turn detection mode */
+  turnDetection?: "server_vad" | "manual";
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/types/realtime.ts` around lines 60 - 68, The JSDoc comments above the
RealtimeSession type fields are misaligned so IDEs will show wrong tooltips;
move each comment to immediately precede its corresponding field (`temperature?:
number;` should have `/** Temperature for AI responses */`, `instructions?:
string;` should have `/** Instructions/system prompt for the session */`,
`turnDetection?: "server_vad" | "manual";` should have `/** Turn detection mode
*/`, and `tools?: RealtimeTool[];` should have `/** Tools/functions available to
the model */`) so the comments attach to the correct symbols (`temperature`,
`instructions`, `turnDetection`, `tools`, and type `RealtimeTool`).
🧹 Nitpick comments (7)
src/lib/voice/providers/OpenAITTS.ts (1)

101-108: getVoices language filter is a no-op.

Both branches return OpenAITTS.VOICES unchanged, so languageCode has no effect. Either drop the parameter/branch, or actually filter (e.g. return [] or a warning for unsupported languages). The current shape misleads callers into thinking language filtering happens.

♻️ Proposed simplification
-  async getVoices(languageCode?: string): Promise<TTSVoice[]> {
-    // OpenAI voices are pre-defined, filter by language if provided
-    if (languageCode && !languageCode.startsWith("en")) {
-      // OpenAI TTS works with multiple languages but voices are English-named
-      return OpenAITTS.VOICES;
-    }
+  async getVoices(_languageCode?: string): Promise<TTSVoice[]> {
+    // OpenAI voices are multilingual; the language parameter does not affect
+    // voice selection, so the full list is always returned.
     return OpenAITTS.VOICES;
   }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/OpenAITTS.ts` around lines 101 - 108, The getVoices
method in OpenAITTS (function getVoices) currently ignores languageCode because
both branches return OpenAITTS.VOICES; fix by either removing the languageCode
parameter/branch entirely or implement real filtering: inspect OpenAITTS.VOICES
for a language or locale property and return only matching entries, or if voices
truly only support English return an empty array (or a logged warning) when
languageCode is present and doesn't start with "en" (e.g. if languageCode &&
!languageCode.startsWith("en") return []). Update the logic in getVoices and
adjust any callers if you remove the parameter.
src/cli/factories/sagemakerCommandFactory.ts (1)

242-259: tempClient is constructed but never exercised — truthy check is dead.

new SageMakerClient(...) either throws or returns a non-null object, so !tempClient is always false. The block validates nothing beyond what the subsequent !secureConfig.accessKeyId/!secureConfig.secretAccessKey checks already cover. Either drop the client instantiation or actually perform a lightweight call (e.g. ListEndpointsCommand with a small page) to validate credentials, since the comment claims "this will throw if credentials are invalid" but no network call is made.

♻️ Proposed simplification
   private static validateSecureConfiguration(
     secureConfig: SecureConfiguration,
   ): void {
-    // Create temporary AWS SDK client with secure credentials
-    const tempClient = new SageMakerClient({
-      region: secureConfig.region,
-      credentials: {
-        accessKeyId: secureConfig.accessKeyId,
-        secretAccessKey: secureConfig.secretAccessKey,
-      },
-    });
-
-    // Test basic connectivity (this will throw if credentials are invalid)
-    // Note: We're not actually making a call here, just validating the client can be created
-    if (
-      !tempClient ||
-      !secureConfig.accessKeyId ||
-      !secureConfig.secretAccessKey
-    ) {
+    if (!secureConfig.accessKeyId || !secureConfig.secretAccessKey) {
       throw new Error("Invalid AWS credentials provided");
     }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/cli/factories/sagemakerCommandFactory.ts` around lines 242 - 259, The
tempClient instantiation (SageMakerClient) is not actually validating
credentials because the truthy check on tempClient is useless; replace the dead
check by performing a lightweight API call to validate credentials (or remove
the client creation if you prefer no runtime validation). Concretely, after
creating tempClient call await tempClient.send(new
ListEndpointsCommand({MaxResults: 1})) (or another cheap call), catch any thrown
error and convert it into the existing throw new Error("Invalid AWS credentials
provided") (including the original error message for debugging); ensure you
import ListEndpointsCommand, await the promise, and only treat missing
secureConfig.accessKeyId/secretAccessKey as a pre-check before calling the API.
src/lib/server/voice/voiceWebSocketHandler.ts (1)

241-290: Consider guarding against runaway reconnects and timer re-entrancy.

Two small resilience gaps in the Soniox plumbing:

  1. connectSoniox() retries every 500ms with no backoff or attempt cap. If Soniox is down or returns a permanent auth/quota error, each session will hammer the endpoint indefinitely and spam logs until the client disconnects.
  2. startKeepAlive() overwrites keepAliveTimer without clearing an existing timer. Today the only call site is inside open, which is paired with stopKeepAlive() on close, so it's safe; but a stray double-invocation would leak the previous interval. A defensive clear is essentially free.
🔧 Proposed tweak
-    function startKeepAlive() {
-      keepAliveTimer = setInterval(() => {
+    function startKeepAlive() {
+      stopKeepAlive();
+      keepAliveTimer = setInterval(() => {
         if (sonioxWs?.readyState === WebSocket.OPEN) {
           sonioxWs.send(JSON.stringify({ type: "keepalive" }));
         }
       }, 8000);
     }

For backoff, consider tracking a reconnectAttempts counter per session and using Math.min(500 * 2 ** attempts, 30000) with a reset on successful open.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/server/voice/voiceWebSocketHandler.ts` around lines 241 - 290, Add a
per-session reconnectAttempts counter used by connectSoniox() to stop runaway
reconnects and apply exponential backoff: on ws.close (when !sessionClosed)
schedule reconnect using delay = Math.min(500 * 2**reconnectAttempts, 30000),
increment reconnectAttempts each attempt, and reset reconnectAttempts = 0 in the
ws.on("open") handler; also ensure you check sessionClosed before scheduling.
Defensively handle keepAliveTimer in startKeepAlive()/stopKeepAlive(): clear any
existing keepAliveTimer at the start of startKeepAlive() before creating a new
setInterval and keep the stopKeepAlive() behavior to clear and null it. Use the
existing symbols connectSoniox, startKeepAlive, stopKeepAlive, keepAliveTimer,
sonioxWs, and sessionClosed to locate and apply these changes.
src/lib/core/baseProvider.ts (1)

1037-1054: !ttsProvider branch is unreachable.

ttsProvider = options.tts?.provider ?? options.provider ?? this.providerName. this.providerName is an AIProviderName set in the constructor and is always a non-empty string, so !ttsProvider can never be true and the warn/return path at Line 1040–1054 is dead. If the intent was to only warn when a caller-specified provider was missing, the check should fire before falling back to this.providerName; otherwise the branch can be dropped.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/core/baseProvider.ts` around lines 1037 - 1054, The guard that logs
and returns when !ttsProvider is unreachable because ttsProvider is computed as
options.tts?.provider ?? options.provider ?? this.providerName (and
this.providerName is always set); update the code by either (A) moving the
missing-provider check to run before falling back to this.providerName (e.g.,
inspect options.tts?.provider and options.provider directly and log/return if
both are absent when the caller expected a specific provider), or (B) remove the
!ttsProvider branch entirely and only keep the aiResponse emptiness check;
locate the ttsProvider assignment and the subsequent if (!aiResponse ||
!ttsProvider) block in baseProvider.ts and apply one of these fixes so the
warning path is reachable or eliminated.
src/lib/voice/RealtimeVoiceAPI.ts (1)

380-393: Make realtime cleanup await disconnects.

clearHandlers() starts disconnects and immediately clears state, so tests or callers can proceed while sockets are still closing.

♻️ Proposed direction
-  static clearHandlers(): void {
+  static async clearHandlers(): Promise<void> {
...
-        handler.disconnect().catch(() => {
+        await handler.disconnect().catch(() => {
           // Ignore errors during cleanup
         });
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/RealtimeVoiceAPI.ts` around lines 380 - 393, clearHandlers()
starts async disconnects but clears this.handlers and this.sessions immediately;
change it to await all handler.disconnect() promises before clearing state.
Iterate this.sessions (or this.handlers) to collect promises from
handler.disconnect() (guard with handler?.isConnected()), use Promise.allSettled
to wait for completion and ignore individual errors, then call
this.handlers.clear(), this.sessions.clear(), and logger.debug only after all
disconnects have settled. Ensure you reference the existing clearHandlers(),
this.sessions, this.handlers, and handler.disconnect() symbols when making the
change.
src/lib/voice/providers/DeepgramSTT.ts (1)

491-492: Floating promise from sendAudio().

sendAudio() is invoked without void or a .catch(). Any rejection inside (other than the caught block at Line 484) would become an unhandled rejection. The OpenAIRealtime/Azure/Gladia equivalents all use void sendAudio(); please match.

-    // Start sending audio in background
-    sendAudio();
+    // Start sending audio in background
+    void sendAudio();

Also, (error as Error).message at Line 497 is an unnecessary cast — error is already typed Error | null and the enclosing if (error) has already narrowed it.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/DeepgramSTT.ts` around lines 491 - 492, Change the
floating promise by calling sendAudio() with the void operator (void
sendAudio()) so any rejection is intentionally ignored the same way as other
providers; also remove the unnecessary cast "(error as Error).message" and use
"error.message" directly since error is already typed and narrowed in the
enclosing if block. Ensure changes are applied around the sendAudio invocation
and the error handling in the same function (sendAudio / surrounding scope).
src/lib/voice/providers/OpenAIRealtime.ts (1)

80-108: Connect-phase listeners leak; later errors can cross-fire into a settled promise.

The open/error handlers registered via this.ws!.on(...) at Lines 85 and 90 are never removed after the connection promise settles. After connect() succeeds, a subsequent socket error will still invoke the old reject(err) (no-op on a settled promise) alongside the intended handler at Line 106, and the orphan open handler will remain forever. Use once (or remove them in a finally) so that only the persistent message/close/error handlers installed after connect remain.

♻️ Suggested refactor
-        this.ws!.on("open", () => {
-          clearTimeout(timeout);
-          resolve();
-        });
-
-        this.ws!.on("error", (err) => {
-          clearTimeout(timeout);
-          reject(err);
-        });
+        this.ws!.once("open", () => {
+          clearTimeout(timeout);
+          resolve();
+        });
+        this.ws!.once("error", (err) => {
+          clearTimeout(timeout);
+          reject(err);
+        });

Same pattern applies to waitForSessionCreated (Lines 292–322) — use once or explicit off in both resolve and reject paths (current code only removes on matching events, so timeouts leak the listener).

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/OpenAIRealtime.ts` around lines 80 - 108, The
connect() method registers persistent "open" and "error" handlers with this.ws
using .on which are never removed, causing cross-fire into the settled
connection promise; change those connect-phase handlers to .once (or register
them and explicitly remove them in a finally block) so the promise's
resolve/reject handlers are not leaked, ensure the timeout is always cleared,
and keep the persistent handlers (this.ws.on("message"/"close"/"error")) only
after the promise resolves; apply the same fix in waitForSessionCreated (replace
ephemeral .on listeners with .once or remove them on both resolve and reject) so
timeouts or later socket events do not leak listeners or trigger callbacks on a
settled promise.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 22646063-aeed-476c-bc94-0951b34755ea

📥 Commits

Reviewing files that changed from the base of the PR and between 5bceb89 and f4cd22b.

⛔ Files ignored due to path filters (1)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (46)
  • CHANGELOG.md
  • memory-bank/voice-bridge-implementation-plan.md
  • package.json
  • src/cli/factories/commandFactory.ts
  • src/cli/factories/sagemakerCommandFactory.ts
  • src/cli/loop/optionsSchema.ts
  • src/lib/adapters/stt/assemblyaiSTTHandler.ts
  • src/lib/adapters/stt/azureSTTHandler.ts
  • src/lib/adapters/stt/deepgramSTTHandler.ts
  • src/lib/adapters/stt/gladiaSTTHandler.ts
  • src/lib/adapters/stt/googleSTTHandler.ts
  • src/lib/adapters/stt/whisperSTTHandler.ts
  • src/lib/core/baseProvider.ts
  • src/lib/factories/providerRegistry.ts
  • src/lib/neurolink.ts
  • src/lib/observability/exporters/laminarExporter.ts
  • src/lib/observability/exporters/posthogExporter.ts
  • src/lib/observability/utils/spanSerializer.ts
  • src/lib/server/voice/voiceWebSocketHandler.ts
  • src/lib/types/generate.ts
  • src/lib/types/index.ts
  • src/lib/types/realtime.ts
  • src/lib/types/server.ts
  • src/lib/types/span.ts
  • src/lib/types/stream.ts
  • src/lib/types/stt.ts
  • src/lib/types/tts.ts
  • src/lib/types/voice.ts
  • src/lib/utils/sttProcessor.ts
  • src/lib/voice/RealtimeVoiceAPI.ts
  • src/lib/voice/STTProvider.ts
  • src/lib/voice/audio-utils.ts
  • src/lib/voice/errors.ts
  • src/lib/voice/index.ts
  • src/lib/voice/providers/AzureSTT.ts
  • src/lib/voice/providers/AzureTTS.ts
  • src/lib/voice/providers/DeepgramSTT.ts
  • src/lib/voice/providers/ElevenLabsTTS.ts
  • src/lib/voice/providers/GeminiLive.ts
  • src/lib/voice/providers/GoogleSTT.ts
  • src/lib/voice/providers/GoogleTTS.ts
  • src/lib/voice/providers/OpenAIRealtime.ts
  • src/lib/voice/providers/OpenAISTT.ts
  • src/lib/voice/providers/OpenAITTS.ts
  • src/lib/voice/stream-handler.ts
  • test/continuous-test-suite-voice.ts
💤 Files with no reviewable changes (2)
  • CHANGELOG.md
  • src/lib/types/server.ts

Comment thread package.json
{
"name": "@juspay/neurolink",
"version": "9.55.4",
"version": "9.55.2",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Confirm what versions are already published on npm
curl -s https://registry.npmjs.org/@juspay/neurolink | jq '.["dist-tags"], ([.versions | keys[]] | sort | .[-5:])'

Repository: juspay/neurolink

Length of output: 143


Version number regressed (9.55.4 → 9.55.2) — blocks npm publish.

The PR downgrades the version in package.json, but npm registry already has 9.55.4 published as the latest. Publishing this version will fail because npm rejects versions equal to or lower than what is already on the registry.

🔧 Fix
-  "version": "9.55.2",
+  "version": "9.55.4",

Or bump to the next appropriate version (e.g., 9.56.0 if this contains significant changes).

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"version": "9.55.2",
"version": "9.55.4",
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@package.json` at line 3, The package.json "version" field was regressed from
9.55.4 to 9.55.2 which will block npm publish; update the "version" value in
package.json to a forward-incremented semver (e.g., 9.56.0 or the next
appropriate patch/minor depending on changes) so it is greater than 9.55.4, and
ensure any release notes or CI/version-related checks that reference the
"version" field are consistent with this new value.

Comment thread src/lib/adapters/stt/azureSTTHandler.ts Outdated
Comment on lines +310 to +317
// Create WebSocket connection
const ws = new WebSocket(wsUrl, {
// @ts-expect-error - headers are supported by Node.js WebSocket libraries
headers: {
"Ocp-Apim-Subscription-Key": this.subscriptionKey,
},
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🌐 Web query:

Does Node.js 18/20 built-in global WebSocket (undici) support a headers option in the constructor, or does it silently ignore non-standard options like the ws npm package allows?

💡 Result:

No, the built-in global WebSocket from undici in Node.js 18/20 does not support a headers option directly in the constructor. The constructor accepts new WebSocket(url[, protocols]), where the second argument can be a string, array of strings for subprotocols, or a WebSocketInit object primarily for dispatcher (e.g., custom agent) and protocols. Official Undici documentation and examples do not list headers as a supported option in WebSocketInit. A comment in the undici PR adding WebSocketInit (#2088) mentions it "conveniently lets people set headers for the ws handshake," but this appears Node.js-specific or outdated, as current docs (v6.x+) do not confirm headers support and examples only show dispatcher/protocols. Node.js docs and tutorials show only basic usage without options. Unlike the ws npm package, which explicitly supports { headers: {...} } and silently handles non-standard options, undici's spec-compliant WebSocket does not; passing unsupported options like headers may be ignored or error, but is not documented/supported. To set custom headers, use a custom Dispatcher (e.g., Client with headers), but this requires undici import, not global WebSocket.

Citations:


🏁 Script executed:

# Verify azureSTTHandler.ts at lines 310-317
if [ -f "src/lib/adapters/stt/azureSTTHandler.ts" ]; then
  echo "=== azureSTTHandler.ts lines 310-317 ==="
  sed -n '310,317p' "src/lib/adapters/stt/azureSTTHandler.ts"
  echo ""
fi

# Check DeepgramSTT.ts for the claimed correct pattern at lines 392-397
if [ -f "src/lib/adapters/stt/deepgramSTTHandler.ts" ]; then
  echo "=== deepgramSTTHandler.ts lines 390-400 (context around 392-397) ==="
  sed -n '390,400p' "src/lib/adapters/stt/deepgramSTTHandler.ts"
  echo ""
fi

# Check gladiaSTTHandler.ts at lines 254-259
if [ -f "src/lib/adapters/stt/gladiaSTTHandler.ts" ]; then
  echo "=== gladiaSTTHandler.ts lines 254-259 ==="
  sed -n '254,259p' "src/lib/adapters/stt/gladiaSTTHandler.ts"
  echo ""
fi

Repository: juspay/neurolink

Length of output: 957


🏁 Script executed:

# Check if ws package is imported elsewhere and how event handlers are attached
echo "=== Checking for ws package imports ==="
rg -n "import.*from.*['\"]ws['\"]" src/lib/adapters/stt/

echo ""
echo "=== Checking for dynamic import of ws ==="
rg -n "await import\(.*ws" src/lib/adapters/stt/

echo ""
echo "=== Checking onmessage vs on('message') patterns ==="
rg -n "on\(['\"]message" src/lib/adapters/stt/ | head -20

Repository: juspay/neurolink

Length of output: 196


🏁 Script executed:

# Get line counts to understand file sizes
wc -l src/lib/adapters/stt/azureSTTHandler.ts src/lib/adapters/stt/deepgramSTTHandler.ts src/lib/adapters/stt/gladiaSTTHandler.ts 2>/dev/null || echo "Files not found"

Repository: juspay/neurolink

Length of output: 213


🏁 Script executed:

# Find where WebSocket is created in deepgramSTTHandler.ts
echo "=== Searching for WebSocket instantiation in deepgramSTTHandler.ts ==="
rg -n "new WebSocket\|await import\(.*ws\)" src/lib/adapters/stt/deepgramSTTHandler.ts -A 3 -B 1

echo ""
echo "=== Full deepgramSTTHandler.ts WebSocket-related code ==="
rg -n "WebSocket\|ws\|headers" src/lib/adapters/stt/deepgramSTTHandler.ts -A 2 -B 1 | head -40

Repository: juspay/neurolink

Length of output: 189


🏁 Script executed:

# Check imports at the top of deepgramSTTHandler.ts
echo "=== deepgramSTTHandler.ts imports (first 50 lines) ==="
head -50 src/lib/adapters/stt/deepgramSTTHandler.ts

Repository: juspay/neurolink

Length of output: 1437


🏁 Script executed:

# Check how event handlers are attached in both files
echo "=== Event handlers in azureSTTHandler.ts ==="
rg -n "\.on\(|onmessage|onerror|onopen|onclose" src/lib/adapters/stt/azureSTTHandler.ts | head -10

echo ""
echo "=== Event handlers in deepgramSTTHandler.ts ==="
rg -n "\.on\(|onmessage|onerror|onopen|onclose" src/lib/adapters/stt/deepgramSTTHandler.ts | head -10

Repository: juspay/neurolink

Length of output: 453


🏁 Script executed:

# Find ws variable assignment in deepgramSTTHandler.ts
echo "=== Searching for ws variable in deepgramSTTHandler.ts ==="
rg -n "const ws\s*=|let ws\s*=|var ws\s*=" src/lib/adapters/stt/deepgramSTTHandler.ts -A 5

echo ""
echo "=== Broader search for WebSocket or ws in deepgramSTTHandler.ts ==="
rg -n "ws\." src/lib/adapters/stt/deepgramSTTHandler.ts | head -20

Repository: juspay/neurolink

Length of output: 881


🏁 Script executed:

# Check a wider range around the event handler lines to see context
echo "=== deepgramSTTHandler.ts lines 270-290 ==="
sed -n '270,290p' src/lib/adapters/stt/deepgramSTTHandler.ts

echo ""
echo "=== deepgramSTTHandler.ts lines 300-320 ==="
sed -n '300,320p' src/lib/adapters/stt/deepgramSTTHandler.ts

Repository: juspay/neurolink

Length of output: 1308


All three STT handlers (azure, deepgram, gladia) fail to send auth headers—they use undici's global WebSocket which silently ignores the headers option.

The undici WebSocket (Node 18+) conforms to WHATWG spec and does not support a headers option in the constructor. The Ocp-Apim-Subscription-Key header at lines 310–317 is never sent; Azure closes with 401. The same bug exists at:

  • deepgramSTTHandler.ts lines 270–275 (Authorization header)
  • gladiaSTTHandler.ts lines 254–259 (x-gladia-key header)

All three require dynamic import of the ws npm package, which supports custom headers:

🔧 Fix pattern
-    const ws = new WebSocket(wsUrl, {
-      // `@ts-expect-error` - headers are supported by Node.js WebSocket libraries
+    const { default: NodeWebSocket } = await import("ws");
+    const ws = new NodeWebSocket(wsUrl, {
       headers: {
-        "Ocp-Apim-Subscription-Key": this.subscriptionKey,
+        "Ocp-Apim-Subscription-Key": this.subscriptionKey!,
       },
     });

Event handler syntax differs: ws package uses on("message", ...) instead of onmessage property. Update all four handler assignments in each file.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/azureSTTHandler.ts` around lines 310 - 317, The
WebSocket constructors in azureSTTHandler (ws variable), deepgramSTTHandler, and
gladiaSTTHandler are using the WHATWG/undici global WebSocket which ignores the
headers option; replace the constructor usage by dynamically importing the "ws"
npm package (e.g., const {default: Ws} = await import("ws")) and instantiate new
Ws(wsUrl, { headers: { ... } }) so the Ocp-Apim-Subscription-Key / Authorization
/ x-gladia-key headers are actually sent; also update the event wiring from
onmessage/onopen/onclose/onerror to the "ws" package style ws.on('message',
...), ws.on('open', ...), ws.on('close', ...), ws.on('error', ...) and preserve
existing handler logic in the Azure handler (and the analogous handlers in
deepgramSTTHandler.ts and gladiaSTTHandler.ts).

Comment on lines +269 to +275
// Create WebSocket connection
const ws = new WebSocket(wsUrl, {
// @ts-expect-error - headers are supported by Node.js WebSocket libraries
headers: {
Authorization: `Token ${this.apiKey}`,
},
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🌐 Web query:

Does Deepgram streaming /v1/listen accept authentication via the token query parameter instead of the Authorization header?

💡 Result:

No, Deepgram's streaming /v1/listen endpoint (WebSocket at wss://api.deepgram.com/v1/listen) does not accept authentication via a generic "token" query parameter. Authentication must use the Authorization header with "Token YOUR_API_KEY" or "Bearer YOUR_JWT_TOKEN" (from /v1/auth/grant). For browser/client-side WebSocket connections where custom headers are blocked, use Sec-WebSocket-Protocol: "token, YOUR_API_KEY" instead. Query parameters are for options like model, language, etc., not auth tokens. Official docs confirm header-based auth exclusively, with no mention of token query param support.

Citations:


WebSocket auth headers won't work with WHATWG WebSocket in Node ≥22, but query parameter auth is not the solution.

This handler attempts to pass headers to the WebSocket constructor, but Node.js's built-in WHATWG WebSocket (Node ≥22) and browsers do not accept a headers option in the constructor. The second argument is reserved for subprotocols only, so the Authorization header will be silently dropped at runtime, and Deepgram will reject the connection with 401.

However, Deepgram's /v1/listen endpoint does not accept authentication via a query parameter. Authentication requires the Authorization: Token <API_KEY> header. Per Deepgram's official documentation, for client-side connections where custom headers are unavailable, use the Sec-WebSocket-Protocol header instead:

Sec-WebSocket-Protocol: token, <API_KEY>

Update the WebSocket instantiation to use the subprotocol approach for WHATWG compatibility:

🔧 Suggested fix
-    const ws = new WebSocket(wsUrl, {
-      // `@ts-expect-error` - headers are supported by Node.js WebSocket libraries
-      headers: {
-        Authorization: `Token ${this.apiKey}`,
-      },
-    });
+    const ws = new WebSocket(wsUrl, [`token, ${this.apiKey}`]);
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/deepgramSTTHandler.ts` around lines 269 - 275, The
WebSocket constructor currently passes a headers option (in the ws creation
block using wsUrl and this.apiKey) which WHATWG WebSocket (Node ≥22 / browsers)
ignores; replace that with the Sec-WebSocket-Protocol style auth by passing the
subprotocol(s) as the second argument instead of headers (e.g., supply "token,
<API_KEY>" or an array with that string) when instantiating the WebSocket (the
code that creates const ws = new WebSocket(wsUrl, { headers: { Authorization:
`Token ${this.apiKey}` } }) should be changed to call new WebSocket(wsUrl,
'token, <API_KEY>' or ['token, <API_KEY>']) using this.apiKey, and remove the
headers/ts-expect-error approach so Deepgram receives the token via
Sec-WebSocket-Protocol).

Comment on lines +114 to +117
constructor(apiKey?: string, baseUrl?: string) {
this.apiKey = apiKey ?? process.env.GLADIA_API_KEY ?? null;
this.baseUrl = baseUrl ?? "https://api.gladia.io/v2";
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🌐 Web query:

Gladia STT v2 API real-time streaming WebSocket endpoint URL and initialization flow

💡 Result:

Gladia STT v2 real-time streaming uses WebSocket. There is no direct static WebSocket endpoint; instead, follow this initialization flow: 1. POST to https://api.gladia.io/v2/live with header x-gladia-key: YOUR_API_KEY and JSON body specifying audio config (required: encoding e.g. "wav/pcm", sample_rate e.g. 16000, bit_depth e.g. 16, channels e.g. 1; optional: model e.g. "solaria-1", language_config, etc.). Example curl: curl --request POST \ --url https://api.gladia.io/v2/live \ --header 'Content-Type: application/json' \ --header 'x-gladia-key: YOUR_GLADIA_API_KEY' \ --data '{ "encoding": "wav/pcm", "sample_rate": 16000, "bit_depth": 16, "channels": 1 }' Response: { "id": "uuid", "url": "wss://api.gladia.io/v2/live?token=uuid" } 2. Connect WebSocket to the returned url (wss://api.gladia.io/v2/live?token=SESSION_ID). 3. Send audio chunks as binary frames or JSON {"type": "audio_chunk", "data": {"chunk": "base64audio"}}. 4. Send {"type": "stop_recording"} to end. 5. Receive events like transcripts (type: "transcript"), speech_start, etc. To get final results: GET https://api.gladia.io/v2/live/{id}. Note: This differs from v1 which had a fixed wss://api.gladia.io/audio/text/audio-transcription with config on connect.

Citations:


🏁 Script executed:

cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '230,250p'

Repository: juspay/neurolink

Length of output: 876


🏁 Script executed:

cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '240,280p'

Repository: juspay/neurolink

Length of output: 1528


🏁 Script executed:

cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '280,320p'

Repository: juspay/neurolink

Length of output: 1483


🏁 Script executed:

cat -n src/lib/adapters/stt/gladiaSTTHandler.ts | sed -n '320,360p'

Repository: juspay/neurolink

Length of output: 1565


WebSocket implementation uses v1 API instead of v2.

The code hardcodes the v1 WebSocket endpoint (wss://api.gladia.io/audio/text/audio-transcription, line 240) and sends configuration on connect (lines 351-359), but the constructor defaults baseUrl to https://api.gladia.io/v2 (line 116). Gladia v2 requires a different flow: POST to https://api.gladia.io/v2/live with audio configuration to obtain a session-specific token, then connect to wss://api.gladia.io/v2/live?token=SESSION_ID. The current implementation will fail against the v2 API. Update to use the v2 initialization flow or remove the v2 baseUrl if v1 is intentional.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/stt/gladiaSTTHandler.ts` around lines 114 - 117, The
WebSocket code uses the v1 endpoint and sends config on connect while the
constructor defaults baseUrl to v2; update the Gladia STT adapter to follow the
v2 flow: in the constructor/initialization for GladiaSTTHandler (constructor and
the method that opens the socket), perform a POST to
https://api.gladia.io/v2/live with the audio/config payload to receive the
session token, then open the WebSocket to
wss://api.gladia.io/v2/live?token=SESSION_TOKEN and stop sending the
configuration over the socket on connect; ensure apiKey handling (this.apiKey)
is passed in the POST, and remove or change any hardcoded
wss://api.gladia.io/audio/text/audio-transcription usage so all endpoints match
the v2 flow (or if you intended v1, change the constructor default baseUrl to
the v1 base and keep the existing ws flow).

Comment thread src/lib/voice/audio-utils.ts
Comment thread src/lib/voice/providers/GoogleSTT.ts
Comment on lines +47 to +49
getSupportedFormats(): AudioFormat[] {
return ["opus"]; // OpenAI Realtime uses PCM16/g711 but we expose as opus for simplicity
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Emitted audio chunks are labeled as opus but actual bytes are PCM16 — consumers will fail to decode.

sendSessionUpdate configures output_audio_format: "pcm16" (Line 245) yet getSupportedFormats() advertises "opus" and every emitted RealtimeAudioChunk sets format: "opus" (Lines 340, 352). Any downstream consumer that dispatches decoding based on chunk.format will try to decode raw PCM as Opus. Either request Opus output from OpenAI or label the chunks correctly as pcm16.

🔧 Suggested fix
   getSupportedFormats(): AudioFormat[] {
-    return ["opus"]; // OpenAI Realtime uses PCM16/g711 but we expose as opus for simplicity
+    return ["pcm16"];
   }
-          this.emitAudio({
-            data: audioData,
-            index: this.audioChunkIndex++,
-            isFinal: false,
-            format: "opus",
-            sampleRate: 24000,
-          });
+          this.emitAudio({
+            data: audioData,
+            index: this.audioChunkIndex++,
+            isFinal: false,
+            format: "pcm16",
+            sampleRate: 24000,
+          });

(apply symmetrically to the response.audio.done branch)

Also applies to: 336-354

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/voice/providers/OpenAIRealtime.ts` around lines 47 - 49, The
getSupportedFormats() return value and emitted RealtimeAudioChunk format labels
are incorrect: change getSupportedFormats() to include "pcm16" (or return only
"pcm16") and update every place that sets chunk.format to "opus" in
OpenAIRealtime (including the RealtimeAudioChunk construction in the audio data
handling and the response.audio.done branch) to "pcm16"; alternatively, if you
prefer Opus, modify sendSessionUpdate to request output_audio_format: "opus" and
ensure emitted chunks are actual Opus bytes—make the change consistently in
getSupportedFormats(), sendSessionUpdate, and where chunks are created to keep
format labels and bytes in sync.

Comment thread src/lib/voice/stream-handler.ts
Comment thread src/lib/voice/stream-handler.ts
@github-actions

github-actions Bot commented May 1, 2026 •

Copy link
Copy Markdown
Contributor

Documentation Validation Results

⚠️ Documentation validation has issues

Check Status Result
Frontmatter Validation ✅ Passed
TypeScript Check ✅ Passed
Build ❌ Failed
Link Validation ⏩ Skipped

🚧 Please fix the failing checks before merging.

Commit: 4eb7a8528208da9e312b7b3fdbd0fb409e386b9c | Workflow: View logs

@murdore

murdore commented May 1, 2026

Copy link
Copy Markdown
Contributor Author

Review Feedback Addressed (Cycle 1)

Files Deleted (all comments on these are now moot)

  • src/lib/adapters/stt/ (all 6 handlers) — dead code, wrong interface
  • src/lib/voice/providers/GoogleTTS.ts — duplicate of existing googleTTSHandler
  • src/lib/voice/STTProvider.ts — duplicate STTProcessor
  • src/lib/voice/voiceAgent.ts, compositeVoice.ts, voiceFactory.ts, voiceRegistry.ts — parallel infrastructure removed
  • src/cli/commands/voice.ts — removed
  • src/lib/neurolink.ts — synthesize(), transcribe(), startRealtimeVoice() methods removed (everything goes through generate()/stream() now)

Fixes Applied to Remaining Code

File Fix Review Comment
OpenAIRealtime.ts Fixed call_id — now requires model-provided ID, skips if missing Copilot #2
OpenAIRealtime.ts Audio format already correct (wav for pcm16, not opus) CR #17
AzureSTT.ts supportsStreaming = false already set CR #10
AzureTTS.ts SSML xml:lang attribute now XML-escaped CR Review 2
OpenAITTS.ts Unsupported formats now log warning instead of silent coercion CR Review 1
ElevenLabsTTS.ts languageCode filter already implemented in getVoices() CR Review 1
realtime.ts JSDoc comments realigned with correct fields CR Review 2
GoogleSTT.ts getAccessToken() now passes keyFilename to GoogleAuth CR #16
DeepgramSTT.ts Added 10s WebSocket connection timeout CR Review 2
GeminiLive.ts Function call errors now propagated via emitError() CR Review 1
RealtimeVoiceAPI.ts Event handler detachment on failed connect already implemented CR Review 2
commandFactory.ts TTS format choices expanded to full AudioFormat union CR Review 2
audio-utils.ts splitIntoChunks/resamplePcm guards already implemented CR #15
stream-handler.ts chunkSize validation + backpressure + pendingData already fixed CR #18, #19

STT Pipeline Bug Fixes

  • Transcription now injected into LLM prompt (was computed but never used)
  • Empty text validation skipped when stt.audio is provided
  • Google STT: sampleRateHertz omitted for WAV/FLAC (reads from header)

Validation

  • TypeScript: ✅ (0 errors)
  • Lint: ✅ (0 errors, 28 pre-existing warnings)
  • Build: ✅ (publint clean)
  • All providers tested with real API calls (Google TTS, OpenAI TTS, ElevenLabs, Azure TTS, Whisper, Deepgram, Google STT, Azure STT)

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented May 1, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

@murdore Thanks for the detailed Cycle 1 summary — that's a thorough set of changes. A few highlights worth noting:

  • The architectural cleanup (removing the duplicate STTProvider.ts, dead adapters/stt/ handlers, and parallel voice infrastructure) significantly reduces surface area and potential for confusion.
  • Routing everything through generate()/stream() rather than standalone synthesize()/transcribe() methods on NeuroLink is a cleaner API boundary.
  • The STT-prompt injection fix is a meaningful correctness improvement — the transcription was previously computed but silently discarded.

Triggering a full review now against the updated state of the branch.

[full_review]


🧠 Learnings used
Learnt from: vigneshJuspay
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-12-15T18:35:37.783Z
Learning: In juspay/neurolink TTS implementation (PR `#691`), the new StreamChunk discriminated union type introduced in TTS-019 will be integrated with StreamResult.stream during the actual TTS streaming implementation PR (TTS-020/TTS-021), not in the type-definition PR. This phased approach keeps type updates and implementation changes atomic.

Learnt from: murdore
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2026-04-17T04:30:25.521Z
Learning: In neurolink (src/lib/neurolink.ts), the `pipelineAHandled: true` flag on `generation:end` events is set ONLY by `finalizeGenerateRequestResult` (the SDK-level wrapper). Native providers (amazonBedrock.ts, ollama.ts, googleAiStudio.ts, googleVertex.ts) emit their own `generation:end` events WITHOUT this flag. The `initializeMetricsListeners()` early-return guard `if (data.pipelineAHandled) return` correctly prevents duplicates: it skips the SDK-level event while allowing native provider events through to Pipeline B.

Learnt from: vigneshJuspay
Repo: juspay/neurolink PR: 693
File: src/lib/core/baseProvider.ts:490-517
Timestamp: 2025-12-18T15:13:28.435Z
Learning: In juspay/neurolink TTS integration (PR `#693`), when options.provider is "auto" and passed to TTSProcessor.synthesize, it will fail automatically during handler lookup since only concrete providers ("google-ai", "vertex") are registered as TTS handlers. No explicit validation against "auto" is needed—the implicit failure at handler registration lookup is by design.

Learnt from: vigneshJuspay
Repo: juspay/neurolink PR: 237
File: memory-bank/tts-provider-implementation-plan.md:92-106
Timestamp: 2025-11-17T13:53:20.209Z
Learning: In PR 237's TTS modality implementation approach, TTS functionality uses GOOGLE_AI_API_KEY (not GOOGLE_TTS_API_KEY) when using the google-ai provider. TTS is implemented as an output modality that leverages the existing google-ai provider authentication.

Learnt from: vigneshJuspay
Repo: juspay/neurolink PR: 693
File: src/lib/core/baseProvider.ts:490-517
Timestamp: 2025-12-18T15:13:28.435Z
Learning: In juspay/neurolink TTS architecture (PR `#693`), timeout handling for TTS synthesis is intentionally managed at the provider/handler level (within TTS handler implementations like GoogleTTSHandler), not at the BaseProvider orchestration level. This allows each provider to enforce its own timeout constraints appropriate to its synthesis capabilities.

Learnt from: vigneshJuspay
Repo: juspay/neurolink PR: 693
File: src/lib/utils/ttsProcessor.ts:318-319
Timestamp: 2025-12-18T15:21:37.311Z
Learning: In juspay/neurolink TTS implementation (PR `#693`), timeout handling for TTSProcessor.synthesize() is intentionally delegated to individual TTS handler implementations (e.g., GoogleTTSHandler) rather than enforced at the TTSProcessor orchestration level. This design allows each provider to implement its own timeout constraints appropriate to its API characteristics and requirements.

@murdore
murdore force-pushed the feat/voice-speech-integration branch from ad077b9 to ae33100 Compare May 1, 2026 09:59
@murdore
murdore force-pushed the feat/voice-speech-integration branch from ae33100 to 636e140 Compare May 1, 2026 10:05
@murdore

murdore commented May 1, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented May 1, 2026

Copy link
Copy Markdown
✅ Actions performed

Full review triggered.

@murdore

murdore commented May 1, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented May 1, 2026

Copy link
Copy Markdown
✅ Actions performed

Full review triggered.

Add Speech-to-Text (STT) as a new capability alongside existing TTS,
with multi-provider support for both. Everything flows through the
existing generate() and stream() JSON config pattern.

New TTS providers (via generate({ tts: { provider: "..." } })):
- openai-tts: OpenAI TTS API (tts-1, tts-1-hd), 6 voices
- elevenlabs: ElevenLabs (eleven_multilingual_v2)
- azure-tts: Azure Cognitive Services Speech

New STT providers (via generate({ stt: { enabled: true, audio, provider } })):
- whisper/openai-stt: OpenAI Whisper API
- google-stt: Google Cloud Speech-to-Text
- deepgram: Deepgram Nova-2/Nova-3
- azure-stt: Azure Cognitive Services Speech

Realtime providers (registered for future SDK use):
- openai-realtime: OpenAI Realtime API (WebSocket)
- gemini-live: Google Gemini Live (WebSocket)

Infrastructure:
- STTProcessor (mirrors TTSProcessor) with SpanType.STT observability
- Audio utilities: format detection, WAV creation, PCM resampling
- ChunkedAudioStream with backpressure and validation
- 30-second fetch timeout on all voice provider API calls
- CLI flags: --stt, --stt-provider, --input-audio, --stt-language, --tts-provider

STT pipeline in generate():
- When stt.audio provided without text: transcription becomes the prompt
- When stt.audio provided with text: transcription prepended as context
- result.transcription contains STTResult with text + confidence

Tested end-to-end with real API calls:
- Google TTS, OpenAI TTS, ElevenLabs, Azure TTS (valid MP3 output)
- Whisper STT (0.95), Deepgram (1.0), Google STT (0.98), Azure STT (0.9)
- Full round-trip: Whisper→Vertex LLM→ElevenLabs (audio in → audio out)
@murdore
murdore force-pushed the feat/voice-speech-integration branch from 4878a57 to 1e4be5f Compare May 1, 2026 10:25
@murdore murdore closed this May 1, 2026
@murdore
murdore deleted the feat/voice-speech-integration branch May 1, 2026 10:26
@murdore
murdore restored the feat/voice-speech-integration branch May 1, 2026 10:26
@murdore murdore reopened this May 1, 2026
@murdore

murdore commented May 1, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented May 1, 2026

Copy link
Copy Markdown
✅ Actions performed

Full review triggered.

@murdore

murdore commented May 1, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #1005 (fresh branch to fix CI workflow trigger issue)

@murdore murdore closed this May 1, 2026

This branch was successfully deployed

1 active deployment
Preview — 1e4be5f9 Deployed May 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants