diff --git a/pmoves/docker-compose.media.yml b/pmoves/docker-compose.media.yml index d848e12cc0..f86543e4e7 100644 --- a/pmoves/docker-compose.media.yml +++ b/pmoves/docker-compose.media.yml @@ -301,7 +301,15 @@ services: - NATS_URL=${NATS_URL} - SUPABASE_URL=${SUPABASE_INTERNAL_URL:-${SUPABASE_URL:-http://supabase-kong:8000}} - TENSORZERO_URL=${TENSORZERO_URL:-http://tensorzero-gateway:3000} - - DEFAULT_VOICE_PROVIDER=${DEFAULT_VOICE_PROVIDER:-vibevoice} + # Default TTS = OmniVoice (k2-fsa, Apache-2.0; license-clean creator-pipeline + # voice). Overridable: DEFAULT_VOICE_PROVIDER=vibevoice restores z890's lane. + - DEFAULT_VOICE_PROVIDER=${DEFAULT_VOICE_PROVIDER:-omnivoice} + # OmniVoice server (creator-operator omnivoice.compose.yml, :8002). On a + # shared host bring OmniVoice up with OMNIVOICE_BIND=0.0.0.0 so this container + # reaches it via host.docker.internal; set OMNIVOICE_TOKEN once 0.0.0.0-bound. + - OMNIVOICE_URL=${OMNIVOICE_URL:-http://host.docker.internal:8002} + - OMNIVOICE_TOKEN=${OMNIVOICE_TOKEN:-} + - OMNIVOICE_TIMEOUT_SEC=${OMNIVOICE_TIMEOUT_SEC:-300} # Realtime duplex voice agent (/v1/voice/agent: mic→Whisper→LLM→VibeVoice). # Off by default — requires pipecat-ai installed + VibeVoice (7860) running. - PIPECAT_ENABLED=${PIPECAT_ENABLED:-false} diff --git a/pmoves/docker-compose.yml b/pmoves/docker-compose.yml index 1a4b915590..996484a88f 100644 --- a/pmoves/docker-compose.yml +++ b/pmoves/docker-compose.yml @@ -3545,7 +3545,15 @@ services: - NATS_URL=${NATS_URL} - SUPABASE_URL=${SUPABASE_INTERNAL_URL:-${SUPABASE_URL:-http://supabase-kong:8000}} - TENSORZERO_URL=${TENSORZERO_URL:-http://tensorzero-gateway:3000} - - DEFAULT_VOICE_PROVIDER=${DEFAULT_VOICE_PROVIDER:-vibevoice} + # Default TTS = OmniVoice (k2-fsa, Apache-2.0; license-clean creator-pipeline + # voice). Overridable: DEFAULT_VOICE_PROVIDER=vibevoice restores z890's lane. + - DEFAULT_VOICE_PROVIDER=${DEFAULT_VOICE_PROVIDER:-omnivoice} + # OmniVoice server (creator-operator omnivoice.compose.yml, :8002). On a + # shared host bring OmniVoice up with OMNIVOICE_BIND=0.0.0.0 so this container + # reaches it via host.docker.internal; set OMNIVOICE_TOKEN once 0.0.0.0-bound. + - OMNIVOICE_URL=${OMNIVOICE_URL:-http://host.docker.internal:8002} + - OMNIVOICE_TOKEN=${OMNIVOICE_TOKEN:-} + - OMNIVOICE_TIMEOUT_SEC=${OMNIVOICE_TIMEOUT_SEC:-300} # Realtime duplex voice agent (/v1/voice/agent: mic→Whisper→LLM→VibeVoice). # Off by default — requires pipecat-ai installed + VibeVoice (7860) running. - PIPECAT_ENABLED=${PIPECAT_ENABLED:-false} diff --git a/pmoves/docs/handoffs/flute-omnivoice-wiring-2026-06-26.md b/pmoves/docs/handoffs/flute-omnivoice-wiring-2026-06-26.md new file mode 100644 index 0000000000..8f829227e8 --- /dev/null +++ b/pmoves/docs/handoffs/flute-omnivoice-wiring-2026-06-26.md @@ -0,0 +1,63 @@ +--- +graphiti_mark: handoff.flute-omnivoice-wiring.2026-06-26 +branch: chore/omnivoice-4090-bringup +pr_numbers: [1885] +scope: Wire flute-gateway → OmniVoice — add OMNIVOICE_URL/TOKEN/TIMEOUT to the + flute-gateway env and flip the default TTS provider (env template + compose) to omnivoice. +risks: Flips the fleet default vibevoice→omnivoice (env-reversible via DEFAULT_VOICE_PROVIDER); + touches z890's vibevoice lane (path intact). Joint host deploy needs OMNIVOICE_BIND=0.0.0.0 + token. +next_actions: Joint-deploy bring OmniVoice up 0.0.0.0+OMNIVOICE_TOKEN; add OMNIVOICE_TOKEN to secrets pipeline; z890 awareness of default flip. +chit_artifact_path: n/a (docs + compose wiring; no CHIT artifact) +agent_signature: 4090-claude +--- + +# Flute-Gateway → OmniVoice Wiring (2026-06-26) + +**By:** 4090-CLAUDE. **Known Road handoff** authorizing the compose edit +(`KNOWN_ROAD=compose:handoff:flute-omnivoice-wiring-2026-06-26.md`). Operator +authorized this wiring directly ("wire flute-gateway + run a TTS smoke" → yes). + +## Context + +OmniVoice is **activated + healthy on the 4090** (`pmoves-omnivoice:latest`, +cuda:0, :8002 — see `[[project_omnivoice_activation]]`). Two TTS smokes passed: +- **Direct:** `POST /synthesize {"text":...}` → 215 KB RIFF WAV. +- **Provider:** the flute-gateway `OmniVoiceProvider` (`health_check` + `synthesize`) + drove the live server → 192 KB RIFF WAV. The integration code path is proven. + +The remaining gap was **deployed config**: flute-gateway's env had +`DEFAULT_VOICE_PROVIDER` (default `vibevoice`) but **no `OMNIVOICE_URL`/`OMNIVOICE_TOKEN`**, +so a deployed flute-gateway would hit its own container loopback, not OmniVoice. + +## Edit (source of truth = ROOT `docker-compose.yml`, NOT the overlay) + +`pmoves/docker-compose.media.yml` is **generated** from `pmoves/docker-compose.yml` +by `scripts/split_compose.py` (editing the overlay alone is dropped on regen). +So the flute-gateway env is edited in the root compose, then overlays are +regenerated with `make -C pmoves compose-split` (drift-gated by `compose-split-check` +/ the `ci/compose-split-gate` CI job, #1881). + +flute-gateway `environment:` changes: +- add `OMNIVOICE_URL=${OMNIVOICE_URL:-http://host.docker.internal:8002}` + (host-gateway resolves to the host where OmniVoice listens; flute-gateway + `main.py` already normalizes a 127.0.0.1 OmniVoice URL → host.docker.internal). +- add `OMNIVOICE_TOKEN=${OMNIVOICE_TOKEN:-}` (sent as `X-OmniVoice-Token`). +- add `OMNIVOICE_TIMEOUT_SEC=${OMNIVOICE_TIMEOUT_SEC:-300}`. +- flip `DEFAULT_VOICE_PROVIDER` default `vibevoice` → **`omnivoice`** (the wiring + intent; OmniVoice is the license-clean creator-pipeline TTS). **Still overridable** + — `DEFAULT_VOICE_PROVIDER=vibevoice` restores the prior default. + +## Runtime requirement (for joint deployment — not in this PR) + +When flute-gateway and OmniVoice run on the **same host**, bring OmniVoice up with +`OMNIVOICE_BIND=0.0.0.0` (+ set `OMNIVOICE_TOKEN`) so the flute-gateway container +can reach it via `host.docker.internal:8002` — a `127.0.0.1` bind is host-loopback +only and unreachable from sibling containers on Linux/bridge. The 4090 validation +ran OmniVoice host-only (`127.0.0.1`) deliberately; the provider smoke ran from the +host. Token is required once bound to 0.0.0.0 (tailnet-exposed). + +## Lane note + +`vibevoice` is **z890's** voice lane (`z890-compose-voice-vibevoice-media-profile.md`). +Flipping the *default* to omnivoice is reversible via env and doesn't remove the +vibevoice path — flagged for z890 awareness. See `[[reference_creator_pipeline_models]]`. diff --git a/pmoves/env.shared.example b/pmoves/env.shared.example index c658ad64a6..6f75195b45 100644 --- a/pmoves/env.shared.example +++ b/pmoves/env.shared.example @@ -288,7 +288,9 @@ WHISPER_URL=http://ffmpeg-whisper:8078 # Ultimate TTS Studio - Multi-engine TTS ULTIMATE_TTS_STUDIO_IMAGE=ghcr.io/powerfulmoves/pmoves-ultimate-tts-studio:pmoves-latest ULTIMATE_TTS_STUDIO_HOST_PORT=7861 -DEFAULT_VOICE_PROVIDER=vibevoice +# Default flute-gateway TTS. omnivoice = k2-fsa OmniVoice (Apache-2.0, license-clean +# creator-pipeline voice). Set to vibevoice to use z890's VibeVoice lane instead. +DEFAULT_VOICE_PROVIDER=omnivoice # Flute Gateway - Multimodal voice communication FLUTE_BASE_URL=http://localhost:8055