Skip to content

feat: add audio transcription pipeline with Telegram voice notes - #289

Closed
serrrfirat wants to merge 3 commits into
mainfrom
feat/audio-pipeline-stt
Closed

serrrfirat wants to merge 3 commits into
mainfrom
feat/audio-pipeline-stt

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • Adds speech-to-text transcription pipeline using OpenAI Whisper API, enabling WASM channels to emit audio attachments that are automatically transcribed before reaching the agent
  • Implements end-to-end Telegram voice note support: voice messages are downloaded via Bot API, sent to Whisper for transcription, and delivered as plain text to the agent
  • Extends WIT interface with attachment-kind/attachment records and adds Attachment/AttachmentKind types to the core channel layer

Details

Architecture: WASM channels download audio (they have HTTP + credential injection), host transcribes (centralized, one implementation serves all channels). Transcription happens at the EmittedMessage → IncomingMessage boundary via TranscriptionMiddleware.

New files:

  • src/transcription/mod.rs — TranscriptionProvider trait, AudioFormat enum, TranscriptionMiddleware, TranscriptionError
  • src/transcription/openai.rs — OpenAI Whisper provider (multipart POST, 25MB file limit)
  • src/config/transcription.rs — TranscriptionConfig with env var overrides (TRANSCRIPTION_ENABLED, TRANSCRIPTION_PROVIDER, TRANSCRIPTION_MODEL, TRANSCRIPTION_LANGUAGE)

Modified files:

  • wit/channel.wit — Added attachment-kind enum and attachment record to emitted-message
  • src/channels/channel.rs — Added Attachment, AttachmentKind, attachments on IncomingMessage
  • src/channels/wasm/host.rs — EmittedMessage gains attachments field with 10MB per-attachment size validation
  • src/channels/wasm/wrapper.rs — Bridges WIT attachments → Rust, applies transcription in both process_emitted_messages and dispatch_emitted_messages
  • channels-src/telegram/src/lib.rs — Parses voice field, downloads via getFile + /file/bot, emits with attachment
  • src/settings.rs — Added TranscriptionSettings
  • src/config/mod.rs, src/lib.rs, src/main.rs — Wired transcription module and config

Graceful degradation:

  • Voice download failure → agent sees "[Voice note: download failed]"
  • Transcription failure → agent sees "[Voice note: transcription failed: {error}]"
  • Transcription disabled → agent sees "[Voice note]" (attachment still carried)
  • Oversized attachments → individual attachment dropped, message still delivered

Closes #90

Test plan

  • cargo fmt — clean
  • cargo clippy --all --all-features — clean
  • cargo test — 1379 passed, 0 failed
  • 16 transcription-specific tests pass (AudioFormat, TranscriptionMiddleware with mock, OpenAI provider, config resolution)
  • Manual: Telegram voice note → agent receives transcribed text (tested with real Whisper API)
  • Manual: WAV file transcription via OpenAI Whisper API — perfect accuracy
  • Manual: Telegram voice with transcription disabled → agent sees "[Voice note]"

🤖 Generated with Claude Code

…am voice notes (#90)

Adds speech-to-text support so WASM channels can emit audio attachments
that get automatically transcribed before reaching the agent. Telegram
voice notes are the first integration — downloaded via Bot API and
transcribed via OpenAI Whisper.

- Extend WIT with attachment-kind, attachment records on emitted-message
- Add Attachment/AttachmentKind types to channel.rs and IncomingMessage
- Add TranscriptionProvider trait, AudioFormat enum, TranscriptionMiddleware
- Implement OpenAI Whisper provider (multipart POST, 25MB limit)
- Add TranscriptionConfig + TranscriptionSettings with env var overrides
- Parse Telegram voice messages, download via getFile, emit as attachments
- Apply transcription in both process/dispatch emitted message paths
- Graceful degradation: download failures show "[Voice note: download failed]"
- Validate attachment sizes (10MB max), drop oversized without losing message

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot added scope: channel Channel infrastructure scope: channel/wasm WASM channel runtime scope: config Configuration scope: dependencies Dependency updates size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: experienced 6-19 merged PRs labels Feb 21, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @serrrfirat, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the system's ability to interact with users by introducing audio transcription capabilities. It enables agents to understand and respond to voice messages, starting with full support for Telegram voice notes. This integration bridges the gap between spoken language and text-based agent processing, making the system more accessible and versatile.

Highlights

  • Audio Transcription Pipeline: Implemented a speech-to-text transcription pipeline using the OpenAI Whisper API, allowing WASM channels to process audio attachments and convert them into text before reaching the agent.
  • Telegram Voice Note Support: Added end-to-end support for Telegram voice messages, which are now downloaded via the Bot API, sent to Whisper for transcription, and delivered as plain text to the agent.
  • Channel Attachment Interface: Extended the WebAssembly Interface Type (WIT) with attachment-kind and attachment records, and introduced Attachment and AttachmentKind types to the core channel layer to handle binary attachments.
  • Configurability and Graceful Degradation: Introduced new configuration options for transcription (enabled, provider, model, language) and implemented graceful degradation for voice note processing, handling download failures, transcription errors, and oversized attachments.
Changelog
  • Cargo.toml
    • Added the multipart feature to the reqwest dependency to support file uploads.
  • channels-src/telegram/Cargo.toml
    • Removed the [workspace] declaration.
  • channels-src/telegram/src/lib.rs
    • Added voice field to TelegramMessage struct to parse incoming voice notes.
    • Defined TelegramVoice and TelegramFile structs for deserializing Telegram API responses related to voice messages and file metadata.
    • Implemented download_voice_file function to fetch audio bytes from Telegram's file API.
    • Modified handle_message to detect voice notes, download their audio, and attach them to the emitted message.
    • Updated message content logic to provide placeholders for voice notes during download or transcription failures, or when transcription is disabled.
  • channels-src/telegram/telegram.capabilities.json
    • Extended the HTTP allowlist to include /file/bot for downloading Telegram files.
  • src/channels/channel.rs
    • Introduced AttachmentKind enum to categorize binary attachments (Audio, Image, Document).
    • Defined Attachment struct to represent binary data, MIME type, filename, and duration.
    • Added attachments vector to IncomingMessage to carry binary data.
    • Provided with_attachments builder method for IncomingMessage.
  • src/channels/mod.rs
    • Exported the new Attachment and AttachmentKind types.
  • src/channels/wasm/host.rs
    • Imported the Attachment type.
    • Defined MAX_ATTACHMENT_SIZE constant for individual attachments (10 MB).
    • Added attachments field to EmittedMessage struct.
    • Implemented validation for attachment sizes in emit_message, dropping oversized attachments while allowing the message to proceed.
  • src/channels/wasm/wrapper.rs
    • Imported Attachment and AttachmentKind types.
    • Updated logging for emit_message to include attachment count.
    • Added logic to convert WIT attachment types to Rust Attachment types.
    • Introduced transcription_middleware field to WasmChannel struct.
    • Applied transcription middleware to IncomingMessage instances in both dispatch_emitted_messages and process_emitted_messages before sending them to the agent.
  • src/config/mod.rs
    • Added transcription module to the configuration system.
    • Exported TranscriptionConfig.
    • Integrated TranscriptionConfig into the main Config struct.
    • Resolved TranscriptionConfig during overall configuration loading.
  • src/config/transcription.rs
    • Added TranscriptionConfig struct to manage transcription settings, including enabled, provider, openai_api_key, model, and language.
    • Implemented Default trait for TranscriptionConfig.
    • Provided resolve method to load transcription settings from environment variables and default settings, including validation for supported providers.
  • src/lib.rs
    • Added transcription module to the library exports.
  • src/main.rs
    • Initialized transcription_middleware based on the resolved TranscriptionConfig.
    • Injected the transcription_middleware into WasmChannel instances during setup.
  • src/settings.rs
    • Added TranscriptionSettings struct for configuring transcription options.
    • Integrated TranscriptionSettings into the main Settings struct.
  • src/transcription/mod.rs
    • Added AudioFormat enum to represent supported audio types and provide utility methods for extensions and MIME types.
    • Defined TranscriptionError enum to handle various transcription-related errors.
    • Introduced TranscriptionProvider trait for abstracting speech-to-text services.
    • Implemented TranscriptionMiddleware to process incoming messages, detect audio attachments, and use a TranscriptionProvider to transcribe them, updating message content accordingly.
  • src/transcription/openai.rs
    • Added OpenAiWhisper struct, implementing the TranscriptionProvider trait for OpenAI's Whisper API.
    • Handled multipart form data for audio file uploads to the OpenAI API.
    • Included logic for managing API keys, model selection, and error handling specific to the Whisper service.
  • wit/channel.wit
    • Defined attachment-kind enum (audio, image, document) within the channel-host interface.
    • Added attachment record with fields for kind, mime-type, data, filename, and duration-secs.
    • Included an attachments list field in the emitted-message record to support binary attachments.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a significant and well-architected feature for audio transcription, particularly for Telegram voice notes, demonstrating robust implementation with clear separation of concerns and extensibility through the TranscriptionProvider trait. However, it introduces two significant security issues: a potential Denial of Service (DoS) due to missing timeouts in transcription HTTP requests, and a potential credential leak in the Telegram channel from unsanitized user input. Additionally, there's a suggestion to address code duplication in src/channels/wasm/wrapper.rs for improved maintainability.

Comment thread src/channels/wasm/wrapper.rs Outdated

// Apply transcription middleware if available (may replace content with transcript)
if let Some(middleware) = transcription_middleware {
msg = middleware.process(msg).await;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

The polling loop in dispatch_emitted_messages awaits transcription without a timeout, similar to process_emitted_messages. This creates a potential Denial of Service (DoS) vulnerability, as a slow or hanging response from the transcription provider will block the polling loop indefinitely. Additionally, to improve maintainability and avoid code duplication, consider extracting the shared logic for processing and dispatching emitted messages into a private helper function, as this logic is very similar to process_emitted_messages.

Comment thread src/channels/wasm/wrapper.rs Outdated

// Apply transcription middleware if available (may replace content with transcript)
if let Some(ref middleware) = self.transcription_middleware {
msg = middleware.process(msg).await;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

The transcription middleware is invoked synchronously for every emitted message with an audio attachment, but the underlying HTTP request to the OpenAI Whisper API lacks a timeout. Since this call occurs outside the WASM execution timeout boundary, a hanging response from the provider will cause the host's HTTP worker threads to hang indefinitely. An attacker could exploit this by sending audio files that trigger slow processing, leading to resource exhaustion and a denial of service.

Comment on lines +915 to +918
let get_file_url = format!(
"https://api.telegram.org/bot{{TELEGRAM_BOT_TOKEN}}/getFile?file_id={}",
file_id
);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-medium medium

The file_id is concatenated directly into the URL without encoding. This URL is later processed by the host's credential injector. A malicious user could provide a file_id containing credential placeholders (e.g., {OPENAI_API_KEY}), causing the host to inject sensitive secrets into the URL sent to Telegram, leading to credential exposure.

References
  1. Avoid forwarding broad host credentials (such as GITHUB_TOKEN or GH_TOKEN) directly into sandboxed containers. Use scoped credentials or dedicated service accounts with the minimum required permissions to prevent host credential leakage.

serrrfirat and others added 2 commits February 21, 2026 15:13
…, file_id sanitization

- Add 30s tokio::time::timeout around transcription middleware to prevent
  a slow/hanging Whisper API from blocking the message pipeline (DoS)
- Extract shared EmittedMessage→IncomingMessage conversion into
  convert_emitted_to_incoming() helper, eliminating duplication between
  process_emitted_messages and dispatch_emitted_messages
- Sanitize file_id and file_path in Telegram voice download to reject
  curly braces, preventing credential placeholder injection via malicious
  file_id values like "{OPENAI_API_KEY}"

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Keep main's formatted JSON + setup section, preserve our /file/bot
allowlist entry needed for voice note file downloads.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@ilblackdragon

Copy link
Copy Markdown
Member

Everything from #289 is on main: transcription module, Telegram voice download, WIT attachment types,
store-attachment-data host function, TranscriptionMiddleware, config. No gaps.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: experienced 6-19 merged PRs risk: medium Business logic, config, or moderate-risk modules scope: channel/wasm WASM channel runtime scope: channel Channel infrastructure scope: config Configuration scope: dependencies Dependency updates size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: Audio pipeline (speech-to-text, text-to-speech, voice note handling)

2 participants