Skip to content

test(multimodal): add processor-level audio, video, office and TTS suites - #1257

Merged
murdore merged 1 commit into
releasefrom
test/multimodal-suites
Aug 4, 2026
Merged

murdore merged 1 commit into
releasefrom
test/multimodal-suites

Conversation

@murdore

@murdore murdore commented Aug 2, 2026 •

Copy link
Copy Markdown
Contributor

Closes #477, #483, #485, #487, #491, #493, #495, #496, #498, #499, #502, #518, #527.

13 of the 15 filed test issues. #528 is left open deliberately (see below).

Why these were open

The implementation shipped; the tests didn't.

$ ls test/ | grep -iE 'audio|video|office'
(nothing)

…while AudioProcessor, VideoProcessor, WordProcessor and ExcelProcessor are all live in release.

Two levels — 50 passing tests

Suite Tests Level
test:audio 10 processor
test:video 9 + 1 skip processor
test:office 10 processor
test:tts:unit 9 unit
test:multimodal:sdk 12 + 1 skip SDK — generate() / stream()

The SDK suite is the one that matters for #491, #493, #498, #502 and #510: it drives real files through generate(), stream(), FileDetector.detectAndProcess() and buildMultimodalMessagesArray(), and asserts on content the model could only know by reading the file — FALCON inside a .docx, 4242 in a spreadsheet cell, PELICAN in a mixed audio+spreadsheet call. "The call succeeded" cannot be mistaken for "the document arrived".

Full run output and a per-issue evidence map are attached as comments below.

What the SDK level found

#1258 — stream() drops file content on the Vertex path. Identical input, same provider:

Path Result
generate() "ORCHID"
stream() "there are no documents attached to this conversation"

Confirmed at maxTokens: 1024. Anthropic is correct on both paths. The file is detected on the stream path, so the loss is between detection and the Vertex request. The parity test runs against Anthropic; a skipped marker keeps the Vertex gap visible until #1258 lands.

This is precisely the class of defect processor-level tests cannot see, and it was invisible until these tests existed.

#528 left open

It requires real OpenAI/Google/Azure TTS calls across 6 voices with synthesizeStream(). This key set has no OpenAI credits, so I could not verify it — leaving it open rather than claiming it.

Fixtures are generated, not committed

ffmpeg mints the audio/video; a .docx is a ZIP of XML parts built with adm-zip (a direct dependency), and an .xlsx is written by the same exceljs the processor reads back. Real parsers read real container headers, so synthetic bytes prove nothing — and a generated fixture cannot drift out of sync with the format the way a checked-in binary can. Missing tools or optional dependencies SKIP rather than fail.

Assertions corrected against the code

  • AudioProcessor documents graceful degradation: corrupt and zero-byte input return success with zeroed metadata and codec: "unknown". Two tests asserted rejection and failed — the contract was right.
  • TTSProcessor.synthesize is (text, provider, options); handlers expose isConfigured() and return { buffer, format, size }.
  • buildMultimodalMessagesArray takes (options, provider, model) with a nested input.

Two apparent bugs were investigated and withdrawn: an audioFiles path-vs-Buffer discrepancy that turned out to be Vertex intermittently returning an empty completion (the live helper now retries once), and a maxTokens theory disproved by the same call succeeding at 60 and 512.

Findings recorded, not buried

  1. A 0-byte audio file is indistinguishable from a valid silent recording — both yield Duration: 0:00 | Codec: unknown with nothing marking it degraded. Audio analogue of IMG-010: No Empty Image Handling #293.
  2. durationFormatted is inconsistent — "0:00" from audio, "2s" from video.

Verification

pnpm run check 0 errors / 4,795 files
pnpm run lint 0 errors
all five suites 50 pass · 2 skip · 0 fail

Not wired into CI: no workflow currently runs any continuous suite (test (20) is formatting, linting and build). Separate gap, separate change.

Summary by CodeRabbit

  • Tests
    • Added continuous coverage for audio, video, Office document processing, text-to-speech, and multimodal SDK workflows.
    • Expanded validation for media metadata, file detection, corrupt or empty inputs, document extraction, spreadsheet handling, TTS errors, mixed inputs, and generation/streaming consistency.
    • Added deterministic audio, video, DOCX, and XLSX fixtures with graceful handling when optional tools, codecs, packages, or credentials are unavailable.
    • Added scripts for focused and combined multimodal test suites.

Copilot AI review requested due to automatic review settings August 2, 2026 04:16
@github-actions

github-actions Bot commented Aug 2, 2026 •

Copy link
Copy Markdown
Contributor

✅ Single Commit Policy - COMPLIANT

Status: Policy requirements met • 1 commit • Valid format • Ready for merge

📊 View validation details

📝 Commit Details

  • Hash: bfcf5cb94088069b7864a0b6edb892d00569b514
  • Message: test(multimodal): add processor and SDK-level audio, video, office and TTS suites
  • Author: Sachin Sharma

✅ Validation Results

  • Single commit requirement met
  • No merge commits in branch
  • Semantic commit message format verified
  • Ready for squash merge to release branch

🤖 Automated validation by NeuroLink Single Commit Enforcement

@coderabbitai

coderabbitai Bot commented Aug 2, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Added continuous test suites for audio, video, Office documents, TTS, and the multimodal SDK. Added runtime media and document fixture helpers. Added package scripts for individual suites and combined multimodal testing.

Changes

Multimodal testing

Layer / File(s) Summary
Fixture generation and test commands
test/helpers/mediaFixtures.ts, test/helpers/officeFixtures.ts, package.json
Added ffmpeg-based audio and video fixtures, DOCX and XLSX fixtures, corrupt input generation, dependency checks, and package scripts.
Audio processing coverage
test/continuous-test-suite-audio.ts
Added tests for audio detection, routing, metadata extraction, invalid input degradation, and filename propagation.
Office processing coverage
test/continuous-test-suite-office.ts
Added tests for DOCX and XLSX detection, extraction, worksheet processing, and invalid archive handling.
TTS processor coverage
test/continuous-test-suite-tts-unit.ts
Added unit tests for provider registration, synthesis dispatch, validation, typed errors, handler replacement, and voice retrieval.
Video processing coverage
test/continuous-test-suite-video.ts
Added tests for video detection, metadata probing, audio tracks, WebM processing, keyframe extraction, invalid input degradation, and filename propagation.
Multimodal SDK coverage
test/continuous-test-suite-multimodal-sdk.ts
Added tests for file detection, message construction, live audio, video, and document requests, streaming parity, mixed inputs, conditional skips, and the Vertex streaming regression marker.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant TestSuite
  participant NeuroLinkSDK
  participant Provider
  TestSuite->>NeuroLinkSDK: Submit audio, video, or document input
  NeuroLinkSDK->>Provider: Build multimodal request
  Provider-->>NeuroLinkSDK: Return generated or streamed content
  NeuroLinkSDK-->>TestSuite: Expose response content
Loading

Possibly related PRs

Suggested reviewers: tara-ag

🚥 Pre-merge checks | ✅ 2 | ❌ 3

❌ Failed checks (3 warnings)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The PR adds audio coverage, but #477 lacks the required test path, provider selection, size validation, mocked OpenAI transcription, and confirmed 85% coverage. Add test/unit/audioProcessor.test.ts with all acceptance cases and verify at least 85% AudioProcessor coverage.
Out of Scope Changes check ⚠️ Warning The video, office, TTS, SDK multimodal suites, and related fixtures are outside the scope of linked issue #477. Move unrelated suites and fixtures into separate PRs or link them to their respective issues.
Docstring Coverage ⚠️ Warning Docstring coverage is 57.14% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the added processor-level test suites and matches the main testing changes, although it omits SDK-level multimodal tests.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch test/multimodal-suites

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds previously-missing multimodal test coverage by introducing new continuous-test suites for audio, video, office documents, and TTS (unit/no-API), along with helper utilities to synthesize media/office fixtures at test runtime. It also exposes these suites via new package.json scripts so they can be run individually or as a single multimodal group.

Changes:

  • Added runtime-generated fixture helpers for office formats (DOCX/XLSX) and media formats (audio/video via ffmpeg).
  • Added four new continuous test suites: audio, video, office, and TTS unit (no API).
  • Added test:* scripts for the new suites plus a test:multimodal aggregate script.

Reviewed changes

Copilot reviewed 7 out of 7 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
test/helpers/officeFixtures.ts Adds in-memory DOCX/XLSX fixture generation and optional dependency detection.
test/helpers/mediaFixtures.ts Adds ffmpeg-based audio/video fixture generation and corrupt fixture creation.
test/continuous-test-suite-audio.ts Adds no-API audio detection + AudioProcessor coverage using generated fixtures.
test/continuous-test-suite-video.ts Adds no-API video detection + VideoProcessor probing/keyframe coverage using generated fixtures.
test/continuous-test-suite-office.ts Adds no-API Word/Excel processing coverage and ZIP/error-path guards using generated fixtures.
test/continuous-test-suite-tts-unit.ts Adds no-API/unit coverage for TTSProcessor registry/validation/dispatch behavior.
package.json Adds new scripts to run the added multimodal suites individually and together.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread test/helpers/officeFixtures.ts
Comment thread test/helpers/mediaFixtures.ts
@Tara-ag

Tara-ag commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

Review Summary

Decision: APPROVED ✅

This PR adds comprehensive test coverage for multimodal processing features (audio, video, office documents, TTS unit tests) that previously had no dedicated test suites.

Changes Overview

  • 7 files modified/added (all test files + CI configuration)
  • 50 new tests across 4 test suites
  • 0 production code changes - purely additive test coverage

Test Suites Added

Suite Tests Coverage
test:audio 10 Audio metadata extraction, MIME detection, FileDetector routing, graceful degradation
test:video 10 Video probing, keyframe extraction, muxed audio track, error handling
test:office 9 DOCX/XLSX extraction, multi-sheet support, error paths (corrupt/truncated files)
test:tts:unit 9 TTS processor logic, registry dispatch, error codes

Key Strengths

  1. Comprehensive edge case coverage: Corrupt files, empty buffers, unsupported providers, truncated archives
  2. Proper dependency handling: Tests SKIP when ffmpeg/exceljs/mammoth unavailable rather than failing
  3. Real fixture generation: Uses ffmpeg and exceljs at runtime instead of committing binary assets
  4. Type safety: All imports use proper TypeScript patterns
  5. Documentation: Clear comments explaining test rationale and behavior expectations

Verification

Per PR description:

  • ✅ pnpm run check: 0 errors / 4,795 files
  • ✅ pnpm run lint: 0 errors
  • ✅ test:audio: 10/10 passing
  • ✅ test:video: 9 pass, 1 skip (VP8 codec not available in test environment)
  • ✅ test:office: 10/10 passing
  • ✅ test:tts:unit: 9/9 passing

Impact on Existing Code

None - this is a pure test addition PR with zero modifications to production code or public APIs.


The PR is ready to merge. It addresses all 15 open test issues in the multimodal backlog (#477, #483, #485, #487, #491, #493, #495, #496, #498, #499, #502, #510, #518, #527, #528) as documented in the PR description.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (2)
test/continuous-test-suite-tts-unit.ts (1)

77-85: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Assert that TTSOptions reaches the handler.

The test checks the text argument but does not check options. If TTSProcessor.synthesize() drops or changes provider options, this test still passes. Use a non-empty valid TTSOptions value and assert that the stub receives its expected fields.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/continuous-test-suite-tts-unit.ts` around lines 77 - 85, Update the
“synthesize dispatches to the registered handler” test to pass a non-empty valid
TTSOptions object to TTSProcessor.synthesize and assert the registered handler
receives the expected option fields unchanged, while preserving the existing
text, invocation count, and buffer assertions.
test/continuous-test-suite-audio.ts (1)

1-15: 📐 Maintainability & Code Quality | 🔵 Trivial | 🏗️ Heavy lift

Close the remaining gaps from issue #477.

This suite covers metadata extraction, MIME/extension detection, and degraded-input handling. Issue #477 also requests provider selection across scenarios, size validation, supported-format checks, and transcription tests using a mocked OpenAI client, plus a stated 85% AudioProcessor coverage target. None of these appear in this file.

Do you want me to draft the provider-selection, size-validation, and mocked-OpenAI-transcription tests as a follow-up addition to this suite?

Also applies to: 104-130

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/continuous-test-suite-audio.ts` around lines 1 - 15, Extend the audio
test suite beyond metadata and detection to cover issue `#477`’s remaining
requirements: provider selection across scenarios, size-limit validation,
supported-format checks, and transcription using a mocked OpenAI client. Add
assertions for AudioProcessor behavior and ensure the resulting tests target the
stated 85% AudioProcessor coverage goal.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@test/continuous-test-suite-tts-unit.ts`:
- Around line 177-184: Add a TTSProcessor.getVoices method that accepts a
provider and locale, resolves the registered handler through the processor
dispatch path, and delegates voice retrieval to that handler. Update the
“getVoices reaches the handler” test to call TTSProcessor.getVoices(PROVIDER,
"en-US") instead of invoking getHandler and the handler directly, while
preserving the existing non-empty voice assertion.

In `@test/continuous-test-suite-video.ts`:
- Around line 91-118: Relax the hardcoded "aac" assertion in the real MP4
processing test around videoProcessor.processFile so it only verifies that
metadata.audioCodec is present, matching the existing audio-track coverage
pattern. Leave the fixture-specific AAC literal assertion in the separate
audio-track test unchanged.

---

Nitpick comments:
In `@test/continuous-test-suite-audio.ts`:
- Around line 1-15: Extend the audio test suite beyond metadata and detection to
cover issue `#477`’s remaining requirements: provider selection across scenarios,
size-limit validation, supported-format checks, and transcription using a mocked
OpenAI client. Add assertions for AudioProcessor behavior and ensure the
resulting tests target the stated 85% AudioProcessor coverage goal.

In `@test/continuous-test-suite-tts-unit.ts`:
- Around line 77-85: Update the “synthesize dispatches to the registered
handler” test to pass a non-empty valid TTSOptions object to
TTSProcessor.synthesize and assert the registered handler receives the expected
option fields unchanged, while preserving the existing text, invocation count,
and buffer assertions.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 696a0689-35b9-4647-8369-f4c5d6a0f8a2

📥 Commits

Reviewing files that changed from the base of the PR and between 4c3f3a2 and 9ee52f3.

📒 Files selected for processing (7)
  • package.json
  • test/continuous-test-suite-audio.ts
  • test/continuous-test-suite-office.ts
  • test/continuous-test-suite-tts-unit.ts
  • test/continuous-test-suite-video.ts
  • test/helpers/mediaFixtures.ts
  • test/helpers/officeFixtures.ts

Comment thread test/continuous-test-suite-tts-unit.ts
Comment thread test/continuous-test-suite-video.ts
@murdore
murdore force-pushed the test/multimodal-suites branch from 9ee52f3 to f95f5d4 Compare August 2, 2026 08:27
@murdore murdore changed the title test(multimodal): add the audio, video, office and TTS suites that were never written test(multimodal): add processor-level audio, video, office and TTS suites Aug 2, 2026
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag

Tara-ag commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

💬 MINOR: Duration format inconsistency between video and audio processors

VideoProcessor renders duration as "2s" while AudioProcessor uses "0:00" (m:ss) format. Both feed the same model so the inconsistency may confuse LLM prompts.

Suggestion: Align the duration formatting across sibling processors. Either adopt m:ss for both or 2s style consistently.

See line 139 in test/continuous-test-suite-video.ts where this is noted in a comment.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
test/continuous-test-suite-video.ts (1)

91-118: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Relax the hardcoded "aac" audio-codec assertion.

makeVideoFile() does not pass -c:a for this MP4 fixture, so ffmpeg selects the MP4 muxer's default audio codec. That default depends on the ffmpeg build, so line 117 can fail on setups that do not default to native AAC.

Use the same pattern already applied at line 132: assert(Boolean(result.data.metadata.audioCodec), ...).

🐛 Proposed fix
-  assertEqual(metadata.audioCodec, "aac", "muxed audio codec identified");
+  assert(Boolean(metadata.audioCodec), "muxed audio codec identified");
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/continuous-test-suite-video.ts` around lines 91 - 118, In the real MP4
probe test, replace the hardcoded audioCodec equality assertion with the
existing truthiness-check pattern used near line 132, asserting that
result.data.metadata.audioCodec is present without requiring a specific codec.
🧹 Nitpick comments (1)
test/continuous-test-suite-audio.ts (1)

3-15: 📐 Maintainability & Code Quality | 🔵 Trivial | 🏗️ Heavy lift

Add the missing AUDIO-029 audio coverage.

This suite covers metadata extraction, FileDetector routing, degraded input, and text content, but no tests cover provider selection, file-size validation, or mocked OpenAI transcription. Add those cases to test/continuous-test-suite-audio.ts and update the AUDIO-029/#477 header so the listed coverage matches the suite contents.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@test/continuous-test-suite-audio.ts` around lines 3 - 15, Extend the audio
suite’s existing tests to cover provider selection, file-size validation, and
mocked OpenAI transcription, using the same fixtures and test structure already
present in test/continuous-test-suite-audio.ts. Update the AUDIO-029/#477 header
description so its listed coverage accurately includes these new cases.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Duplicate comments:
In `@test/continuous-test-suite-video.ts`:
- Around line 91-118: In the real MP4 probe test, replace the hardcoded
audioCodec equality assertion with the existing truthiness-check pattern used
near line 132, asserting that result.data.metadata.audioCodec is present without
requiring a specific codec.

---

Nitpick comments:
In `@test/continuous-test-suite-audio.ts`:
- Around line 3-15: Extend the audio suite’s existing tests to cover provider
selection, file-size validation, and mocked OpenAI transcription, using the same
fixtures and test structure already present in
test/continuous-test-suite-audio.ts. Update the AUDIO-029/#477 header
description so its listed coverage accurately includes these new cases.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: e07343ba-38eb-49f9-8523-7e86fbebcd2c

📥 Commits

Reviewing files that changed from the base of the PR and between 9ee52f3 and f95f5d4.

📒 Files selected for processing (7)
  • package.json
  • test/continuous-test-suite-audio.ts
  • test/continuous-test-suite-office.ts
  • test/continuous-test-suite-tts-unit.ts
  • test/continuous-test-suite-video.ts
  • test/helpers/mediaFixtures.ts
  • test/helpers/officeFixtures.ts
🚧 Files skipped from review as they are similar to previous changes (3)
  • package.json
  • test/continuous-test-suite-office.ts
  • test/continuous-test-suite-tts-unit.ts

@Tara-ag

Tara-ag commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

🛡️ Yama Review Verdict: CHANGES_REQUESTED

Severity counts — 🔒 CRITICAL: 0 · ⚠️ MAJOR: 0 · 💡 MINOR: 1

🤖 Yama Review Summary

⚠️ The review loop ended early (time-limit, 19 steps) — this verdict is built strictly from gate-verified findings.

Verified findings (3):

  • MINOR: Video duration format inconsistency between processors — test/continuous-test-suite-video.ts:139
  • SUGGESTION: Missing explicit timeout handling in office tests — test/continuous-test-suite-office.ts:45
  • SUGGESTION: TTS unit test handler mock could be more comprehensive — test/continuous-test-suite-tts-unit.ts:89

Findings behind this verdict

  • 💡 MINOR: Video duration format inconsistency between processors — test/continuous-test-suite-video.ts:139
    VideoProcessor uses seconds in metadata (duration: number) while AudioProcessor uses formatted time strings. This creates inconsistent output formats when both processors generate textContent.
  • 💬 SUGGESTION: Missing explicit timeout handling in office tests — test/continuous-test-suite-office.ts:45
    Office document tests don't explicitly set timeouts for file processing operations. While this may work for typical file sizes, it's not defensive against slow parsing or large files.
  • 💬 SUGGESTION: TTS unit test handler mock could be more comprehensive — test/continuous-test-suite-tts-unit.ts:89
    The TTSProcessor unit test stub handler only mocks basic synthesis behavior. It doesn't cover edge cases like voice selection, getVoices() errors, or text length validation that real handlers would have.

@murdore
murdore force-pushed the test/multimodal-suites branch from a0407da to d65d6af Compare August 2, 2026 20:56
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

murdore added a commit that referenced this pull request Aug 3, 2026
AudioProcessor and VideoProcessor each carried their own private
formatDuration and disagreed on the result: a two-second file rendered as
"0:02" from audio and "2s" from video, and a zero duration as "0:00" versus
"0s". Both strings land in the textContent handed to the model, frequently in
the same request when a video's muxed audio track is described alongside it,
so the mismatch reads as two different facts about one file.

Both now delegate to a shared formatMediaDuration(). The explicit-unit form
wins over the clock form because these strings are read by a language model,
not rendered in a player scrubber: "1m 30s" has one reading, while "1:30" is
ambiguous between 1m30s and 1h30m and needs context that may not be present.

Video previously truncated where audio rounded, so the two also disagreed on
fractional durations; the shared helper rounds, and a 2.6s clip now reads "3s"
on both sides. Non-finite and negative inputs render "0s" rather than a
fabricated number — callers reach this on a failed probe.

Raised in review on #1257. Regression test pins the format and the edge cases
across both processors. Full bugfixes suite: 235 passed, 0 failed.
murdore added a commit that referenced this pull request Aug 3, 2026
AudioProcessor and VideoProcessor each carried their own private
formatDuration and disagreed on the result: a two-second file rendered as
"0:02" from audio and "2s" from video, and a zero duration as "0:00" versus
"0s". Both strings land in the textContent handed to the model, frequently in
the same request when a video's muxed audio track is described alongside it,
so the mismatch reads as two different facts about one file.

Both now delegate to a shared formatMediaDuration(). The explicit-unit form
wins over the clock form because these strings are read by a language model,
not rendered in a player scrubber: "1m 30s" has one reading, while "1:30" is
ambiguous between 1m30s and 1h30m and needs context that may not be present.

Video previously truncated where audio rounded, so the two also disagreed on
fractional durations; the shared helper rounds, and a 2.6s clip now reads "3s"
on both sides. Non-finite and negative inputs render "0s" rather than a
fabricated number — callers reach this on a failed probe.

Raised in review on #1257. Regression test pins the format and the edge cases
across both processors. Full bugfixes suite: 235 passed, 0 failed.
@murdore
murdore force-pushed the test/multimodal-suites branch from d65d6af to 3979392 Compare August 3, 2026 05:25
@github-actions

github-actions Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag

Tara-ag commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Review Summary

This PR adds comprehensive processor-level test suites for multimodal functionality (audio, video, office documents, TTS) that previously had no dedicated test coverage. The implementation shipped; the tests didn't.

Files Changed (8 total)

  • package.json - Added 6 new test scripts for audio, office, tts:unit, video, multimodal suites
  • test/continuous-test-suite-audio.ts - Tests AudioProcessor with real MP3/WAV/FLAC fixtures
  • test/continuous-test-suite-multimodal-sdk.ts - SDK integration tests verifying files reach models via generate()/stream()
  • test/continuous-test-suite-office.ts - Tests WordProcessor (mammoth) and ExcelProcessor (exceljs)
  • test/continuous-test-suite-tts-unit.ts - Unit tests for TTSProcessor registry/dispatch logic
  • test/continuous-test-suite-video.ts - Tests VideoProcessor metadata extraction and keyframes
  • test/helpers/mediaFixtures.ts - Runtime fixture generation using ffmpeg
  • test/helpers/officeFixtures.ts - Runtime DOCX/XLSX fixture generation

Review Findings

No issues found. All changes are test-only and follow NeuroLink's established patterns:

  • Tests skip gracefully when optional dependencies unavailable (ffmpeg, exceljs, mammoth, credentials)
  • Real fixtures minted at test time rather than committed (avoids repo bloat)
  • Error paths covered (corrupt files, empty buffers, non-ZIP payloads)
  • No hardcoded secrets or credentials
  • Follows existing testing architecture (tsx-based continuous suites)

Impact on Existing Code

  • Blast radius: ~110 additional files affected (tests calling production processors)
  • No breaking changes to public API
  • Test additions only; zero production code modifications
  • All tests are self-contained and idempotent

Architectural Compliance

✅ All CLAUDE.md rules respected (this is a test-only PR)
✅ No static provider imports outside registry
✅ No type definition file violations
✅ No interface declarations (uses type aliases throughout)
✅ Unique type names with domain prefixes

Decision

APPROVED - This PR adds essential test coverage for multimodal processors that was missing. The tests are well-structured, cover edge cases, and follow NeuroLink's testing conventions. No action required beyond merging.

…d TTS suites

Audio, video, Office document and TTS processing shipped without dedicated
suites. These add processor-level coverage plus an SDK-level suite that proves
a file handed to generate()/stream() actually reaches the model — the seam
where multimodal support really breaks, and one the processor tests cannot
reach.

Every live assertion is written so a refusal cannot satisfy it. That
constraint came from evidence, not caution: an earlier revision asserted the
response contained "VIDEO", which passes on the refusal "No video is
attached.", and a third asserted "RECEIVED" after a prompt that instructed the
model to say it. Assertions now key on values obtainable only from the file —
a filename, a fixture's real 320x240 resolution, a cell value, a duration —
each confirmed against a live negative control that receives no file.

Review round 2 hardened two more of the same shape:

- The Buffer-input audio test asserted only the absence of a sentinel. A live
  negative control answers "I apologize for the confusion…" with no file at
  all, which satisfies that. It now asks for the duration of a deliberately
  odd 7s fixture. Sample rate was tried first and rejected: the model returns
  "44100" with no file attached, so a guessable fact is not evidence.
- The mixed-multimodal test asserted "AUDIO", which the refusal "no audio file
  is attached" contains. It now asserts the filename.

generateNonEmpty() also retries NeuroLink's turn-ended messages, which are
non-empty plausible prose carrying no answer and so pass an emptiness check.

Fixtures are minted at run time with ffmpeg and adm-zip/exceljs rather than
committed. Suites skip when a tool or optional dependency is absent; a
dependency that is present but broken fails loudly instead.

Local run: 49 passed, 4 documented skips, 0 failures. The skips are open
product bugs (#1258, #1259), not missing coverage.

Closes #483
Closes #485
Closes #487
Closes #495
Closes #499
Closes #518
Closes #526
Closes #530
@murdore
murdore force-pushed the test/multimodal-suites branch from 3979392 to bfcf5cb Compare August 4, 2026 14:07
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@Tara-ag Tara-ag left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Summary

Decision: APPROVED ✅

This pull request adds comprehensive test coverage for multimodal file processing (audio, video, office documents, and TTS). All 8 changed files are test files only with zero production code changes.

Findings Summary

  • Total findings: 0
  • Critical/Major issues: None
  • Security concerns: None
  • Breaking changes: None

Review Scope

  • ✅ Reviewed all 8 changed files systematically
  • ✅ Verified no hardcoded secrets or credentials in test files
  • ✅ Confirmed proper error handling and graceful degradation (Skip when dependencies unavailable)
  • ✅ Validated TypeScript types and imports follow project conventions
  • ✅ No breaking changes to SDK API (purely additive tests)

Impact on existing code

  • Changed files: 8 (all test files)
  • Affected flows: None (test-only changes)
  • Production impact: None
  • Risk level: Low - these are test additions that don't affect runtime behavior

Files Reviewed

  1. package.json - Added new test scripts (harmless configuration change)
  2. test/continuous-test-suite-audio.ts - Audio processor tests
  3. test/continuous-test-suite-multimodal-sdk.ts - End-to-end SDK tests
  4. test/continuous-test-suite-office.ts - Word/Excel document tests
  5. test/continuous-test-suite-tts-unit.ts - TTS unit tests
  6. test/continuous-test-suite-video.ts - Video processor tests
  7. test/helpers/mediaFixtures.ts - Media fixture generation helpers
  8. test/helpers/officeFixtures.ts - Office document fixture generation helpers

Notes

  • The VideoProcessor duration format note ("2s" vs "0:00") is a documentation comment acknowledging future work - not a code issue
  • All tests properly skip when ffmpeg or optional dependencies are unavailable
  • Following established patterns from CONTRIBUTING.md (no vitest runner, tsx-based suites)

No inline comments required - this is a clean addition of test coverage with no code quality or security issues found.

@Tara-ag

Tara-ag commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Review Summary for PR #1257

Decision: APPROVED ✅

This pull request adds comprehensive test suites for multimodal features (audio, video, office documents, TTS). All changes are confined to test files and configuration; no production code is modified.

Findings

  • 0 issues found - All tests follow established patterns in the codebase
  • No CRITICAL or MAJOR issues - Test logic is sound, error handling is appropriate
  • Test coverage verified - Comprehensive coverage of new features including edge cases

Files Changed

  1. package.json - Added test scripts for new suites (no functional impact)
  2. test/continuous-test-suite-audio.ts - Audio file support tests
  3. test/continuous-test-suite-multimodal-sdk.ts - SDK-level multimodal integration tests
  4. test/continuous-test-suite-office.ts - Word/Excel document processing tests
  5. test/continuous-test-suite-tts-unit.ts - Unit tests for TTSProcessor
  6. test/continuous-test-suite-video.ts - Video file support tests
  7. test/helpers/mediaFixtures.ts - Utility functions for generating media fixtures
  8. test/helpers/officeFixtures.ts - Utility functions for generating Office fixtures

Impact on Existing Code

  • Blast radius: None - all changes are in the test directory only
  • Affected flows: None - no production code modified
  • Breaking changes: None - purely additive test coverage

Verification Against Project Standards

✅ No hardcoded secrets or credentials
✅ No breaking API changes
✅ Follows TypeScript best practices
✅ Proper error handling throughout
✅ Tests skip gracefully when dependencies unavailable (ffmpeg, optional packages)
✅ Cleanup code included (temp directory removal)
✅ Consistent with existing test suite patterns in the repository

Notes

  • Some utility fixture functions (makeAudioFile, makeVideoFile, makeDocx, makeXlsx) don't have dedicated unit tests, but they're tested indirectly through the main suites
  • Known duration format inconsistency between AudioProcessor ("0:00") and VideoProcessor ("2s") was acknowledged in previous review and is not introduced by this PR
  • All live provider tests include proper timeout guards and skip mechanisms

Review completed following Yama methodology: file-by-file analysis, evidence-based findings, and impact assessment via code knowledge graph.

@murdore
murdore merged commit da9feac into release Aug 4, 2026
16 checks passed
@murdore
murdore deleted the test/multimodal-suites branch August 4, 2026 17:03
@github-actions

github-actions Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 10.8.13 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

AUDIO-029: Create Unit Tests for AudioProcessor

3 participants