Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
117 changes: 117 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -299,6 +299,122 @@ curl -X POST http://localhost:${PORT}/bots \
}'
```

You can attach external prompt context and MCP servers per bot. External
context is loaded before the Pipecat process starts and capped by
`prompt_data_token_limit` using an approximate token budget. URL sources block
localhost and private-network targets by default; set
`PROMPT_DATA_ALLOW_PRIVATE_URLS=true` only in trusted deployments.

You can also select the bot LLM per request with `llm_provider` and
`llm_model`. Supported providers are `openai`, `anthropic`, and `zai`.
Provider credentials and base URLs are server-side environment variables only:
`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `ZAI_API_KEY`, and optional
`ZAI_BASE_URL`. OpenAI defaults to Pipecat's Responses API service for newest
models; set `OPENAI_API_SURFACE=chat` to use the older Chat Completions bridge.
The Z.ai integration uses the OpenAI-compatible Chat Completions bridge;
Anthropic uses Pipecat's native Claude bridge. The API rejects unconfigured
providers before creating the upstream MeetingBaaS bot.

MCP servers are live-query capable only when their config is connectable.
`http`, `streamable_http`, and `sse` servers require `transport`, `url`, and
optional `headers`. Local process `stdio` MCP is intentionally not accepted by
the public API because it would execute caller-supplied commands. Remote MCP
URLs also block localhost and private-network targets by default; production
deployments should use `MCP_ALLOWED_PRIVATE_URLS` with exact `http://host:port/path`
entries for trusted loopback MCPs instead of the broad
`MCP_ALLOW_PRIVATE_URLS=true` development bypass. If `transport` is omitted, the
server is treated as metadata-only and MCP tools are not executed. Secrets are
not required in the request, but `headers` are available for deployments that
need them. Use `tool_allowlist` to constrain which server tools the bot may
call, and set `enabled: false` to document a server without connecting to it.

```bash
curl -X POST http://localhost:${PORT}/bots \
-H "Content-Type: application/json" \
-H "x-meeting-baas-api-key: your-api-key" \
-d '{
"meeting_url": "https://meet.google.com/xxx-yyyy-zzz",
"personas": ["account_executive"],
"llm_provider": "anthropic",
"llm_model": "claude-opus-4-8",
"prompt_data_token_limit": 4000,
"prompt_data_sources": [
{
"name": "CRM account notes",
"type": "url",
"url": "https://example.com/account-notes.md"
},
{
"name": "Call objective",
"type": "text",
"text": "Confirm timeline, budget, and integration constraints."
}
],
"mcp": {
"instructions": "Use Google Drive and CRM context only when relevant.",
"servers": [
{
"name": "drive-docs-work",
"enabled": true,
"transport": "streamable_http",
"url": "http://127.0.0.1:8123/mcp",
"tool_allowlist": ["search", "read"],
"timeout_seconds": 20,
"instructions": "Use only meeting-relevant Drive docs."
},
{
"name": "remote-crm",
"enabled": true,
"transport": "streamable_http",
"url": "https://mcp.example.com/mcp",
"headers": {
"Authorization": "Bearer optional-token"
},
"tools": ["get_account", "list_recent_calls"],
"tool_allowlist": ["get_account", "list_recent_calls"],
"timeout_seconds": 15
}
]
},
"speech_speed": 1.25
}'
```

`speech_speed` overrides `CARTESIA_TTS_SPEED`, `TTS_SPEED`, or
`SPEECH_SPEED`. The Cartesia runner clamps speed to `0.6..1.5`.

LLM defaults are resolved as request value, then provider-specific env, then
generic `LLM_MODEL`, then service default:

- OpenAI: `OPENAI_MODEL`, default `gpt-5.5`. `OPENAI_API_SURFACE` defaults to
`responses`; set it to `chat` for compatibility with older OpenAI-compatible
paths. `OPENAI_SERVICE_TIER` is passed through when set.
- Anthropic: `ANTHROPIC_MODEL`, default `claude-opus-4-8`. Low-latency example:
`claude-haiku-4-5`.
- Z.ai: `ZAI_MODEL`, default `glm-5.2`; `ZAI_BASE_URL` defaults to
`https://api.z.ai/api/paas/v4/`.

### OpenAPI Snapshots

This repo contains three OpenAPI files with different roles:

- `openapi.json` is this FastAPI service snapshot for generic tooling.
- `speaking-bot-openapi.json` is the same service snapshot, named explicitly
for the speaking-bots MCP sync.
- `meeting-baas-openapi-v1.json` is the upstream MeetingBaaS v1 API snapshot.
- `openapi-v2.json` is the upstream MeetingBaaS v2 API snapshot.

The service snapshots include `/bots`, `/bots/{bot_id}`,
`/personas/generate-image`, `/health`, `/ready`, `/webhook`, and the current
`BotRequest` fields for `prompt_data_sources`, `prompt_data_token_limit`, `mcp`,
and `speech_speed`.
Comment on lines +402 to +405

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Snapshot field list omits llm_provider/llm_model.

The "current BotRequest fields" list mentions prompt_data_sources, prompt_data_token_limit, mcp, and speech_speed, but not llm_provider/llm_model, even though LLM provider routing is this cohort's headline feature and both fields are present in the generated schema.

📝 Proposed fix
 The service snapshots include `/bots`, `/bots/{bot_id}`,
 `/personas/generate-image`, `/health`, `/ready`, `/webhook`, and the current
-`BotRequest` fields for `prompt_data_sources`, `prompt_data_token_limit`, `mcp`,
-and `speech_speed`.
+`BotRequest` fields for `prompt_data_sources`, `prompt_data_token_limit`, `mcp`,
+`llm_provider`, `llm_model`, and `speech_speed`.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
The service snapshots include `/bots`, `/bots/{bot_id}`,
`/personas/generate-image`, `/health`, `/ready`, `/webhook`, and the current
`BotRequest` fields for `prompt_data_sources`, `prompt_data_token_limit`, `mcp`,
and `speech_speed`.
The service snapshots include `/bots`, `/bots/{bot_id}`,
`/personas/generate-image`, `/health`, `/ready`, `/webhook`, and the current
`BotRequest` fields for `prompt_data_sources`, `prompt_data_token_limit`, `mcp`,
`llm_provider`, `llm_model`, and `speech_speed`.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@README.md` around lines 402 - 405, The snapshot field list in the README is
missing the `llm_provider` and `llm_model` `BotRequest` fields, so update the
documentation where the current `BotRequest` fields are described to include
both of those schema fields alongside `prompt_data_sources`,
`prompt_data_token_limit`, `mcp`, and `speech_speed`. Keep the existing service
snapshot list unchanged and ensure the wording around `BotRequest` matches the
generated schema used by the snapshot docs.


Regenerate the service snapshot after API model changes:

```bash
poetry run python scripts/export_openapi.py
```

You can still manually specify a WebSocket URL if needed:

```bash
Expand Down Expand Up @@ -440,6 +556,7 @@ Once the server is running, you can access:

- Interactive API docs: `http://localhost:${PORT}/docs`
- OpenAPI specification: `http://localhost:${PORT}/openapi.json`
- Committed service OpenAPI snapshots: `openapi.json`, `speaking-bot-openapi.json`
- Health endpoint: `http://localhost:${PORT}/health`
- Readiness endpoint: `http://localhost:${PORT}/ready`

Expand Down
204 changes: 202 additions & 2 deletions app/models.py
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
"""Data models for the Speaking Meeting Bot API."""

from datetime import datetime
from typing import Any, Dict, List, Optional
from typing import Any, Dict, List, Literal, Optional

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Use built-in collection types and | None annotations.

Ruff flags typing.Dict/typing.List, and the project guideline requires modern built-in annotations. Apply this consistently across the new public models.

Proposed cleanup
-from typing import Any, Dict, List, Literal, Optional
+from typing import Any, Literal

-    text: Optional[str] = Field(
+    text: str | None = Field(
...
-    headers: Optional[Dict[str, str]] = Field(
+    headers: dict[str, str] | None = Field(
...
-    tools: Optional[List[str]] = Field(
+    tools: list[str] | None = Field(
...
-    prompt_data_sources: Optional[List[PromptDataSource]] = Field(
+    prompt_data_sources: list[PromptDataSource] | None = Field(

Also applies to: 67-79, 123-151, 187-192, 248-290

🧰 Tools
🪛 Ruff (0.15.20)

[warning] 4-4: typing.Dict is deprecated, use dict instead

(UP035)


[warning] 4-4: typing.List is deprecated, use list instead

(UP035)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@app/models.py` at line 4, Update the new public models in app/models.py to
use modern built-in collection annotations instead of typing.Dict/typing.List,
and replace Optional with the PEP 604 form using | None. Apply this consistently
across the affected model definitions and any related type hints in the
referenced model blocks, while keeping Literal and Any where appropriate. Focus
on the imports and the model classes/functions that currently use Dict, List, or
Optional so the annotations match the project’s style and Ruff expectations.

Sources: Coding guidelines, Linters/SAST tools


from pydantic import BaseModel, ConfigDict, Field, field_validator
from pydantic import BaseModel, ConfigDict, Field, field_validator, model_validator


def _validate_meeting_url(value: str) -> str:
Expand Down Expand Up @@ -49,6 +49,153 @@ class TurnConfig(BaseModel):
)


class PromptDataSource(BaseModel):
"""External context to append to the bot prompt under a token budget."""

model_config = ConfigDict(extra="forbid")

name: str = Field(
"external_context",
min_length=1,
max_length=120,
description="Human-readable source name shown inside the prompt context block",
)
type: Literal["text", "url"] = Field(
...,
description="Whether to load inline text or fetch an external HTTP(S) URL",
)
text: Optional[str] = Field(
None,
description="Inline context. Required when type is text.",
)
Comment on lines +67 to +70

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

git ls-files app/models.py && printf '\n---\n' && sed -n '1,140p' app/models.py && printf '\n---\n' && rg -n "prompt_data_token_limit|max_length|request-size|body size|content-length|token_limit|inline context|type is text" -S .

Repository: Meeting-BaaS/speaking-meeting-bot

Length of output: 9434


🏁 Script executed:

sed -n '240,320p' app/models.py && printf '\n---\n' && sed -n '1,260p' app/services/prompt_context.py && printf '\n---\n' && sed -n '200,260p' app/routes.py && printf '\n---\n' && rg -n "max_request|body_size|max_body|limit.*bytes|Content-Length|request size|request_size|client_max_body_size|limit_bytes" -S app README.md .

Repository: Meeting-BaaS/speaking-meeting-bot

Length of output: 14926


🏁 Script executed:

rg -n "FastAPI\\(|uvicorn|nginx|client_max_body_size|max_body|request size|body size|Content-Length|limit.*bytes|prompt_data_token_limit|PROMPT_DATA_SOURCE_MAX_BYTES" -S . && printf '\n---\n' && git ls-files | rg '(^|/)(main|app|server|run|docker|nginx|compose|proxy|deploy|README|docs).*'

Repository: Meeting-BaaS/speaking-meeting-bot

Length of output: 7548


Cap inline prompt text app/models.py:67-70 — text is unbounded here, and there’s no request-body limit elsewhere in this path, so a client can send a very large inline payload and pay the parse/validation cost before prompt_data_token_limit applies. Add a max_length here or enforce an upstream body-size cap.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@app/models.py` around lines 67 - 70, The inline prompt field in the model
schema is currently unbounded, so very large requests can be accepted before
`prompt_data_token_limit` is checked. Add a `max_length` constraint to the
`text` field in `app.models` (the `Field` definition for `text`) or enforce an
upstream request-body size limit so oversized inline payloads are rejected
earlier.

url: Optional[str] = Field(
None,
description="HTTP(S) URL to fetch. Required when type is url.",
)
headers: Optional[Dict[str, str]] = Field(
None,
description="Optional HTTP headers for URL sources. Avoid request-specific secrets unless needed.",
)
token_limit: Optional[int] = Field(
None,
ge=1,
le=50_000,
description="Optional per-source token cap before the request-level cap is applied",
)

@field_validator("url")
@classmethod
def validate_url(cls, value: Optional[str]) -> Optional[str]:
if value is None:
return value
normalized = value.strip()
if not normalized.startswith(("http://", "https://")):
raise ValueError("prompt data source url must start with http:// or https://")
Comment on lines +92 to +93

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Do not allow cleartext HTTP for user-configured prompt and MCP URLs.

These URLs can be fetched with caller-supplied headers or MCP credentials, so accepting http:// risks transmitting secrets and prompt data in cleartext. Prefer requiring https://, with any local-dev exception gated explicitly.

Proposed hardening
-        if not normalized.startswith(("http://", "https://")):
-            raise ValueError("prompt data source url must start with http:// or https://")
+        if not normalized.startswith("https://"):
+            raise ValueError("prompt data source url must start with https://")
...
-        if not normalized.startswith(("http://", "https://")):
-            raise ValueError("mcp server url must start with http:// or https://")
+        if not normalized.startswith("https://"):
+            raise ValueError("mcp server url must start with https://")

Also applies to: 163-164

🧰 Tools
🪛 Ruff (0.15.20)

[warning] 93-93: Avoid specifying long messages outside the exception class

(TRY003)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@app/models.py` around lines 92 - 93, The URL validation in the prompt/MCP
model helpers currently allows both HTTP and HTTPS, which should be tightened to
reject cleartext URLs. Update the validation logic in the prompt data source and
MCP URL checks (the helper that normalizes and validates URLs in app/models.py)
so user-configured endpoints require https:// by default, and only allow http://
behind an explicit local-development exception flag or equivalent gate. Ensure
the existing ValueError message and any related validation paths consistently
reflect the new HTTPS-only requirement.

Source: Linters/SAST tools

return normalized

@model_validator(mode="after")
def validate_source_payload(self):
if self.type == "text" and not self.text:
raise ValueError("text is required when prompt data source type is text")
if self.type == "url" and not self.url:
raise ValueError("url is required when prompt data source type is url")
if self.type == "text" and self.url:
raise ValueError("url is not allowed when prompt data source type is text")
if self.type == "url" and self.text:
raise ValueError("text is not allowed when prompt data source type is url")
return self


MCPTransport = Literal["http", "streamable_http", "sse"]
LLMProvider = Literal["openai", "anthropic", "zai"]


class MCPServerConfig(BaseModel):
"""MCP server metadata and optional live connection details."""

model_config = ConfigDict(extra="forbid")

name: str = Field(..., min_length=1, max_length=120)
enabled: bool = Field(
True,
description="Whether this server may be used. Disabled servers are documented but not connected.",
)
url: Optional[str] = Field(
None,
description="Remote MCP server URL. Required for http, streamable_http, and sse transports.",
)
headers: Optional[Dict[str, str]] = Field(
None,
description="Optional HTTP headers for remote MCP servers. Use only when a server requires them.",
)
transport: Optional[MCPTransport] = Field(
None,
description="Remote MCP transport. Omit for metadata-only servers that cannot execute tools.",
)
tools: Optional[List[str]] = Field(
None,
max_length=50,
description="Known tool names exposed by this MCP server",
)
tool_allowlist: Optional[List[str]] = Field(
None,
max_length=50,
description="Optional allowlist of MCP tool names this bot may call from this server.",
)
timeout_seconds: Optional[float] = Field(
None,
ge=0.1,
le=300.0,
description="Optional per-server connection/tool timeout in seconds.",
)
instructions: Optional[str] = Field(
None,
max_length=4_000,
description="Operator instructions or constraints for this MCP server",
)

@field_validator("url")
@classmethod
def validate_mcp_url(cls, value: Optional[str]) -> Optional[str]:
if value is None:
return value
normalized = value.strip()
if not normalized.startswith(("http://", "https://")):
raise ValueError("mcp server url must start with http:// or https://")
return normalized

@model_validator(mode="after")
def validate_connection_details(self):
if self.transport in {"http", "streamable_http", "sse"}:
if not self.url:
raise ValueError(
f"url is required when MCP transport is {self.transport}"
)
else:
if self.url or self.headers:
raise ValueError(
"transport is required when MCP connection details are supplied"
)
return self


class MCPConfig(BaseModel):
"""MCP server metadata and optional live connection details."""

model_config = ConfigDict(extra="forbid")

servers: List[MCPServerConfig] = Field(
default_factory=list,
max_length=10,
description="MCP servers to document and optionally connect for tool calls",
)
instructions: Optional[str] = Field(
None,
max_length=4_000,
description="Global MCP usage instructions for the bot",
)


class BotRequest(BaseModel):
"""Request model for creating a speaking bot in a meeting."""

Expand All @@ -64,6 +211,28 @@ class BotRequest(BaseModel):
"enable_tools": True,
"extra": {"company": "ACME Corp", "meeting_purpose": "Weekly sync"},
"websocket_url": "wss://bots.example.com",
"prompt_data_token_limit": 3000,
"llm_provider": "anthropic",
"llm_model": "claude-opus-4-8",
"prompt_data_sources": [
{
"name": "CRM account notes",
"type": "url",
"url": "https://example.com/account-notes.md",
}
],
"speech_speed": 1.15,
"mcp": {
"servers": [
{
"name": "crm",
"url": "https://mcp.example.com",
"transport": "streamable_http",
"tools": ["get_account", "list_recent_calls"],
"tool_allowlist": ["get_account", "list_recent_calls"],
}
]
},
"prompt": "You are Meeting Assistant, a concise and professional \
AI bot that helps summarize key points and keep the meeting on track. Speak clearly and stay on topic.",
}
Expand Down Expand Up @@ -93,6 +262,37 @@ class BotRequest(BaseModel):
None,
description="Per-bot turn-taking tuning (VAD confidence/start_secs/stop_secs/min_volume)",
)
prompt_data_sources: Optional[List[PromptDataSource]] = Field(
None,
max_length=10,
description="External text or URL data sources to append to the bot prompt",
)
prompt_data_token_limit: int = Field(
4_000,
ge=0,
le=50_000,
description="Approximate total token cap for loaded prompt_data_sources. 0 disables loading.",
)
mcp: Optional[MCPConfig] = Field(
None,
description="MCP server/tool metadata and optional live connection details",
)
llm_provider: Optional[LLMProvider] = Field(
None,
description="LLM provider for this bot. Defaults to LLM_PROVIDER, then openai.",
)
llm_model: Optional[str] = Field(
None,
min_length=1,
max_length=120,
description="Provider model for this bot. Defaults to provider-specific env vars.",
)
speech_speed: Optional[float] = Field(
None,
ge=0.5,
le=2.0,
description="TTS speaking speed multiplier. Defaults to CARTESIA_TTS_SPEED, TTS_SPEED, SPEECH_SPEED, or the runner default.",
)

# NOTE: streaming_audio_frequency is intentionally excluded and handled internally

Expand Down
Loading