Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions docs-site/static/search-index.json

Large diffs are not rendered by default.

2 changes: 2 additions & 0 deletions docs/api/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -835,6 +835,7 @@ console.log(result.content);
- [CliServeFlatRoute](type-aliases/CliServeFlatRoute.md)
- [CliToolRoutingFlags](type-aliases/CliToolRoutingFlags.md)
- [CliClassifierRouterFlags](type-aliases/CliClassifierRouterFlags.md)
- [CliVoicesCommandArgs](type-aliases/CliVoicesCommandArgs.md)
- [CliAgentCommandArgs](type-aliases/CliAgentCommandArgs.md)
- [CliNetworkCommandArgs](type-aliases/CliNetworkCommandArgs.md)
- [CliAudioPlayerCommand](type-aliases/CliAudioPlayerCommand.md)
Expand Down Expand Up @@ -3069,6 +3070,7 @@ console.log(result.content);
- [~~Gender~~](type-aliases/Gender.md)
- [TTSVoice](type-aliases/TTSVoice.md)
- [GoogleAudioEncoding](type-aliases/GoogleAudioEncoding.md)
- [GoogleStreamingAudioEncoding](type-aliases/GoogleStreamingAudioEncoding.md)
- [TTSChunk](type-aliases/TTSChunk.md)
- [CartesiaMessage](type-aliases/CartesiaMessage.md)
- [LogLevel](type-aliases/LogLevel.md)
Expand Down
39 changes: 39 additions & 0 deletions docs/api/classes/GoogleTTSHandler.md
Original file line number Diff line number Diff line change
Expand Up @@ -120,3 +120,42 @@ Audio buffer with metadata
#### Implementation of

`TTSHandler.synthesize`

---

### synthesizeStream()

> **synthesizeStream**(`text`, `options?`): `AsyncIterable`\<[`TTSChunk`](../type-aliases/TTSChunk.md), `any`, `any`\> \| `undefined`

Stream one pre-validated segment's audio as Google produces it.

Returns `undefined` — the contract's "not incrementally deliverable"
signal — unless the voice and the format are both ones the streaming
endpoint was measured to accept, and the text is not SSML.
`StreamingSynthesisInput` has no `ssml` field at all, so markup that
`synthesize()` would honour has to stay on the buffered path rather than
be sent as literal text.

Every non-empty response is yielded as it arrives and carries `isFinal:
false`. `TTSProcessor` recomputes indexes, cumulative sizes and finality
globally across segments and discards whatever a handler reports, so
labelling the last response here would buy nothing and would cost a
one-response lookahead.

#### Parameters

##### text

`string`

##### options?

[`TTSOptions`](../type-aliases/TTSOptions.md) = `{}`

#### Returns

`AsyncIterable`\<[`TTSChunk`](../type-aliases/TTSChunk.md), `any`, `any`\> \| `undefined`

#### Implementation of

`TTSHandler.synthesizeStream`
53 changes: 53 additions & 0 deletions docs/api/classes/TTSProcessor.md
Original file line number Diff line number Diff line change
Expand Up @@ -156,6 +156,59 @@ if (TTSProcessor.supports("google-ai")) {

---

### getVoices()

> `static` **getVoices**(`providerName`, `options?`): `Promise`\<[`TTSVoice`](../type-aliases/TTSVoice.md)[]\>

List the voices a registered provider offers.

The counterpart to `synthesize()` for discovery: a caller cannot pass
`TTSOptions.voice` without first knowing what the provider will accept,
and `getVoices` is optional on `TTSHandler`, so asking the handler
directly means every caller re-implements the same two guards. Both
failures are reported as typed `TTSError`s rather than a `TypeError` on
an absent member.

`languageCode` is passed through verbatim; each handler decides what
filtering it means. Google and Azure query their APIs with it, OpenAI's
voice list is fixed and ignores it.

#### Parameters

##### providerName

`string`

Provider identifier, resolved case-insensitively

##### options?

Optional language filter

###### languageCode?

`string`

#### Returns

`Promise`\<[`TTSVoice`](../type-aliases/TTSVoice.md)[]\>

The provider's voices

#### Throws

TTSError if the provider is not registered or cannot list voices

#### Example

```typescript
const voices = await TTSProcessor.getVoices("google-ai", {
languageCode: "en-US",
});
```

---

### synthesize()

> `static` **synthesize**(`text`, `provider`, `options`): `Promise`\<[`TTSResult`](../type-aliases/TTSResult.md)\>
Expand Down
35 changes: 35 additions & 0 deletions docs/api/type-aliases/CliVoicesCommandArgs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
[**NeuroLink API Reference**](../README.md)

---

[NeuroLink API Reference](../README.md) / CliVoicesCommandArgs

# Type Alias: CliVoicesCommandArgs

> **CliVoicesCommandArgs** = `object`

`neurolink voices` arguments — TTS voice discovery.

## Properties

### provider

> **provider**: `string`

TTS provider whose voices to list (e.g. google-ai, openai-tts).

---

### language?

> `optional` **language?**: `string`

Optional language filter passed through to the provider (e.g. en-US).

---

### json?

> `optional` **json?**: `boolean`

Emit the raw list as JSON instead of a table.
17 changes: 17 additions & 0 deletions docs/api/type-aliases/GoogleStreamingAudioEncoding.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
[**NeuroLink API Reference**](../README.md)

---

[NeuroLink API Reference](../README.md) / GoogleStreamingAudioEncoding

# Type Alias: GoogleStreamingAudioEncoding

> **GoogleStreamingAudioEncoding** = `"PCM"` \| `"OGG_OPUS"`

Audio encodings Google's _streaming_ synthesis endpoint accepts.

Deliberately a separate type from [GoogleAudioEncoding](GoogleAudioEncoding.md): the
batch and streaming endpoints do not accept the same set. `MP3` and
`LINEAR16` are valid for batch and rejected by streaming with
`INVALID_ARGUMENT: Unsupported audio encoding`, while `PCM` — raw
16-bit signed LE, headerless — exists only on the streaming side.
77 changes: 73 additions & 4 deletions docs/features/tts.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,7 +144,36 @@ Google Cloud TTS offers three voice quality tiers:

### Voice Discovery

Voice identifiers follow Google Cloud TTS naming conventions: `<language>-<variant>-<type>-<name>` (e.g., `en-US-Neural2-C`, `en-GB-Wavenet-D`).
`--tts-voice` takes a provider-specific identifier, and `neurolink voices`
prints the ones a provider will accept:

```bash
neurolink voices --provider openai-tts # the six OpenAI voices
neurolink voices --provider google-ai --language en-US # filter by language
neurolink voices --provider elevenlabs --json # machine-readable
```

`--provider` is required. The list is sorted by name and the count is printed
at the end. Providers are registered only when their credentials are present,
so an unconfigured provider is reported the same way a misspelled one is —
with the set that _is_ registered named in the message, and a non-zero exit.

The same list is available from the SDK:

```typescript
import { TTSProcessor } from "@juspay/neurolink";

const voices = await TTSProcessor.getVoices("google-ai", {
languageCode: "en-US",
});
```

`languageCode` is passed through to the provider, which decides what filtering
it means: Google and Azure query their APIs with it, and OpenAI's fixed list
ignores it. An unregistered provider, or one whose handler does not implement
voice listing, raises a `TTS_PROVIDER_NOT_SUPPORTED` error.

Google voice identifiers follow Google Cloud TTS naming conventions: `<language>-<variant>-<type>-<name>` (e.g., `en-US-Neural2-C`, `en-GB-Wavenet-D`).

Refer to the [Google Cloud TTS voice list](https://cloud.google.com/text-to-speech/docs/voices) for all available voices.

Expand Down Expand Up @@ -463,6 +492,14 @@ neurolink generate "Your text" \
# --tts-use-ai-response : synthesize AI response instead of input text
```

**Discovering voice ids for `--tts-voice`:**

```bash
neurolink voices --provider <provider> [--language <code>] [--json]
```

See [Voice Discovery](#voice-discovery).

**Selecting a specific TTS provider:**

```bash
Expand Down Expand Up @@ -656,11 +693,43 @@ use the 3,000-character default.
Handlers can optionally expose provider-native audio reads for each buffered
text segment. NeuroLink prefers that capability when it supports the requested
options and otherwise keeps the existing one-buffer-per-segment synthesis path.
OpenAI TTS currently streams response-body reads for `mp3` and raw `pcm16`;
`wav`, `flac`, `ogg`/`opus`, and other requested formats use buffered synthesis
because native delivery has not been verified for those container formats.
Custom and built-in handlers without the optional capability remain compatible.

Two handlers currently offer it, each only for the options whose incremental
delivery has been verified against the live API:

| Handler | Streams natively for | Everything else |
| ------------ | --------------------------------------------------------------------------------------- | ------------------ |
| `openai-tts` | `mp3`, `pcm16` | buffered synthesis |
| `google-ai` | `pcm16` and `ogg`/`opus`, **and** only for a `Chirp3-HD`, `Chirp-HD` or `Journey` voice | buffered synthesis |

Google's restriction is the streaming endpoint's own, not NeuroLink's: it
rejects every other voice family (`Neural2`, `Studio`, `Wavenet`, `Standard`)
with `only Chirp 3: HD voices are supported for streaming synthesis`, and
rejects `MP3` and `LINEAR16` as unsupported encodings even though
`synthesizeSpeech` accepts both. SSML is also excluded — the streaming request
has no SSML field at all, so `<speak>` input stays on the buffered path rather
than being sent as literal text. Because the default format is `mp3` and the
default voice is a `Neural2` one, native streaming is strictly opt-in: pass
both a streaming-capable voice and `format: "pcm16"` (or `"ogg"`) to get it.

```typescript
// Sentence-buffered audio arrives in many reads instead of one per sentence.
const streamResult = await neurolink.stream({
input: { text: "Summarize the quarterly report." },
tts: {
enabled: true,
provider: "google-ai",
voice: "en-US-Chirp3-HD-Aoede",
format: "pcm16",
},
});
```

Note that `pcm16` is headerless: the chunks report `sampleRate: 24000` so a
consumer can wrap them, and Google's buffered path does not support `pcm16` at
all, so that format only works for the streaming-capable voices above.

Provider-local chunk indexes and finality are not exposed directly. NeuroLink
recomputes a single global zero-based index, cumulative byte size, and exactly
one final chunk across all successful segments. Empty transport reads on the
Expand Down
Loading
Loading