feat: add audio transcription pipeline with Telegram voice notes - #289
serrrfirat wants to merge 3 commits into
Conversation
…am voice notes (#90) Adds speech-to-text support so WASM channels can emit audio attachments that get automatically transcribed before reaching the agent. Telegram voice notes are the first integration — downloaded via Bot API and transcribed via OpenAI Whisper. - Extend WIT with attachment-kind, attachment records on emitted-message - Add Attachment/AttachmentKind types to channel.rs and IncomingMessage - Add TranscriptionProvider trait, AudioFormat enum, TranscriptionMiddleware - Implement OpenAI Whisper provider (multipart POST, 25MB limit) - Add TranscriptionConfig + TranscriptionSettings with env var overrides - Parse Telegram voice messages, download via getFile, emit as attachments - Apply transcription in both process/dispatch emitted message paths - Graceful degradation: download failures show "[Voice note: download failed]" - Validate attachment sizes (10MB max), drop oversized without losing message Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Summary of ChangesHello @serrrfirat, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly enhances the system's ability to interact with users by introducing audio transcription capabilities. It enables agents to understand and respond to voice messages, starting with full support for Telegram voice notes. This integration bridges the gap between spoken language and text-based agent processing, making the system more accessible and versatile. Highlights
Changelog
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces a significant and well-architected feature for audio transcription, particularly for Telegram voice notes, demonstrating robust implementation with clear separation of concerns and extensibility through the TranscriptionProvider trait. However, it introduces two significant security issues: a potential Denial of Service (DoS) due to missing timeouts in transcription HTTP requests, and a potential credential leak in the Telegram channel from unsanitized user input. Additionally, there's a suggestion to address code duplication in src/channels/wasm/wrapper.rs for improved maintainability.
|
|
||
| // Apply transcription middleware if available (may replace content with transcript) | ||
| if let Some(middleware) = transcription_middleware { | ||
| msg = middleware.process(msg).await; |
There was a problem hiding this comment.
The polling loop in dispatch_emitted_messages awaits transcription without a timeout, similar to process_emitted_messages. This creates a potential Denial of Service (DoS) vulnerability, as a slow or hanging response from the transcription provider will block the polling loop indefinitely. Additionally, to improve maintainability and avoid code duplication, consider extracting the shared logic for processing and dispatching emitted messages into a private helper function, as this logic is very similar to process_emitted_messages.
|
|
||
| // Apply transcription middleware if available (may replace content with transcript) | ||
| if let Some(ref middleware) = self.transcription_middleware { | ||
| msg = middleware.process(msg).await; |
There was a problem hiding this comment.
The transcription middleware is invoked synchronously for every emitted message with an audio attachment, but the underlying HTTP request to the OpenAI Whisper API lacks a timeout. Since this call occurs outside the WASM execution timeout boundary, a hanging response from the provider will cause the host's HTTP worker threads to hang indefinitely. An attacker could exploit this by sending audio files that trigger slow processing, leading to resource exhaustion and a denial of service.
| let get_file_url = format!( | ||
| "https://api.telegram.org/bot{{TELEGRAM_BOT_TOKEN}}/getFile?file_id={}", | ||
| file_id | ||
| ); |
There was a problem hiding this comment.
The file_id is concatenated directly into the URL without encoding. This URL is later processed by the host's credential injector. A malicious user could provide a file_id containing credential placeholders (e.g., {OPENAI_API_KEY}), causing the host to inject sensitive secrets into the URL sent to Telegram, leading to credential exposure.
References
- Avoid forwarding broad host credentials (such as GITHUB_TOKEN or GH_TOKEN) directly into sandboxed containers. Use scoped credentials or dedicated service accounts with the minimum required permissions to prevent host credential leakage.
…, file_id sanitization
- Add 30s tokio::time::timeout around transcription middleware to prevent
a slow/hanging Whisper API from blocking the message pipeline (DoS)
- Extract shared EmittedMessage→IncomingMessage conversion into
convert_emitted_to_incoming() helper, eliminating duplication between
process_emitted_messages and dispatch_emitted_messages
- Sanitize file_id and file_path in Telegram voice download to reject
curly braces, preventing credential placeholder injection via malicious
file_id values like "{OPENAI_API_KEY}"
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Keep main's formatted JSON + setup section, preserve our /file/bot allowlist entry needed for voice note file downloads. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
Everything from #289 is on main: transcription module, Telegram voice download, WIT attachment types, |
Summary
attachment-kind/attachmentrecords and addsAttachment/AttachmentKindtypes to the core channel layerDetails
Architecture: WASM channels download audio (they have HTTP + credential injection), host transcribes (centralized, one implementation serves all channels). Transcription happens at the
EmittedMessage → IncomingMessageboundary viaTranscriptionMiddleware.New files:
src/transcription/mod.rs—TranscriptionProvidertrait,AudioFormatenum,TranscriptionMiddleware,TranscriptionErrorsrc/transcription/openai.rs— OpenAI Whisper provider (multipart POST, 25MB file limit)src/config/transcription.rs—TranscriptionConfigwith env var overrides (TRANSCRIPTION_ENABLED,TRANSCRIPTION_PROVIDER,TRANSCRIPTION_MODEL,TRANSCRIPTION_LANGUAGE)Modified files:
wit/channel.wit— Addedattachment-kindenum andattachmentrecord toemitted-messagesrc/channels/channel.rs— AddedAttachment,AttachmentKind,attachmentsonIncomingMessagesrc/channels/wasm/host.rs—EmittedMessagegainsattachmentsfield with 10MB per-attachment size validationsrc/channels/wasm/wrapper.rs— Bridges WIT attachments → Rust, applies transcription in bothprocess_emitted_messagesanddispatch_emitted_messageschannels-src/telegram/src/lib.rs— Parsesvoicefield, downloads viagetFile+/file/bot, emits with attachmentsrc/settings.rs— AddedTranscriptionSettingssrc/config/mod.rs,src/lib.rs,src/main.rs— Wired transcription module and configGraceful degradation:
"[Voice note: download failed]""[Voice note: transcription failed: {error}]""[Voice note]"(attachment still carried)Closes #90
Test plan
cargo fmt— cleancargo clippy --all --all-features— cleancargo test— 1379 passed, 0 failed"[Voice note]"🤖 Generated with Claude Code