Skip to content

feat(video-analysis): add video-analysis support in neurolink - #824

Merged
murdore merged 1 commit into
juspay:releasefrom
rishikagudla:BZ-48424-add-support-for-video-analysis-in-neurolink
Feb 17, 2026
Merged

murdore merged 1 commit into
juspay:releasefrom
rishikagudla:BZ-48424-add-support-for-video-analysis-in-neurolink

Conversation

@rishikagudla

@rishikagudla rishikagudla commented Feb 17, 2026 •

Copy link
Copy Markdown
Contributor

Pull Request

Description

What does this PR do?

Implements comprehensive video analysis support for NeuroLink using Gemini 2.0 Flash. This feature provides deep logical auditing of video sequences, focusing on the "Action-Reaction Chain" to identify silent failures, UI/UX bugs, and logical inconsistencies in recorded workflows.

Related Issues

Does this PR close any issues?

N/A - New feature implementation

Type of Change

Please select the type of change:

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Documentation update
  • Refactoring (no functional changes)
  • Performance improvement
  • Test coverage improvement
  • Build/CI configuration
  • Other (please describe):

Motivation and Context

Why is this change needed? What problem does it solve?

Developers need to analyze screen recordings to debug UI/UX issues, identify workflow failures, and understand logical inconsistencies. Traditional video analysis focuses on visual description, but this implementation provides:

  • Logical Reasoning: Understands "why" and "how" events occur, not just "what" is visible
  • Developer-Focused Reports: Actionable insights structured for engineering teams to fix bugs in one shot
  • Silent Failure Detection: Identifies cases where UI provides no feedback for user actions
  • Action-Reaction Chain Analysis: Traces the logical bond between user intent and system response

Use Cases:

  • Debugging UI/UX flows from screen recordings
  • Identifying where users get stuck in multi-step workflows
  • Detecting silent failures (missing loading states, error feedback)
  • Performance analysis (delays between user actions and system responses)
  • Error handling assessment (clarity of error messages)

Changes Made

What specific changes were made?

Core Implementation

  • Added src/lib/adapters/video/videoAnalyzer.ts with Gemini 2.0 Flash integration for both Vertex AI and Google AI providers
  • Implemented src/lib/utils/videoAnalysisProcessor.ts with frame detection and analysis pipeline orchestration
  • Enhanced src/lib/processors/media/VideoProcessor.ts with adaptive frame extraction using FFmpeg
  • Integrated video analysis into src/lib/core/baseProvider.ts generation flow

Frame Extraction Engine

  • Implemented intelligent temporal sampling: 1 fps for short videos (<10s), adaptive intervals for longer videos
  • Fixed critical FFmpeg bug where the same first frame was extracted 10 times
  • Added prev_selected_t logic in FFmpeg select filter for precise temporal frame distribution
  • Frames are extracted as JPEG at 85% quality for optimal balance between quality and performance

Prompting System

  • Created "Critical Logic Auditor" persona focused on logical reasoning over visual description
  • Structured output into 4 sections: Strategic Overview, Action-Reaction Chain, Critical Findings, Final Assessment
  • Removed all markdown formatting (#, **) to ensure clean terminal output
  • Added operational constraints to prevent model from requesting "more frames" via tools

Documentation

  • Completely rewrote docs/features/video-analysis.md to match the new logic-first workflow
  • Added 4 use case examples: UI/UX bugs, silent failures, workflow validation, comparison analysis
  • Provided CLI and SDK usage examples with advanced configurations (buffers, custom models, disabling tools)
  • Added "Command Gallery" section with quick CLI recipes

Examples

  • Rewrote examples/video-analysis.ts with 6 comprehensive examples:
    1. Standard Logic Audit
    2. Silent Failure Detection
    3. UI/UX Bug Detection
    4. Workflow Validation
    5. Performance & Responsiveness Audit
    6. Error Handling Assessment

Breaking Changes

Does this PR introduce breaking changes?

  • No breaking changes
  • Yes, breaking changes (describe below)

This is an additive feature. Existing functionality remains unchanged. All video analysis is opt-in via providing video files in the input.

Testing

How has this been tested?

  • Manual testing completed
  • Unit tests added/updated - Pending
  • Integration tests added/updated - Pending
  • E2E tests pass
  • Tested with multiple providers: Vertex AI, Google AI Studio
  • Tested on multiple platforms: macOS

Test Coverage

  • All new code is covered by tests - To be added
  • Existing tests pass
  • Coverage percentage maintained or improved - To be verified

Manual Testing Steps

  1. Set up environment with credentials:

    export GOOGLE_VERTEX_PROJECT="your-project-id"
    export GOOGLE_AI_API_KEY="your-api-key"
  2. Test with example file:

    npx tsx examples/video-analysis.ts
  3. Verify that all 6 examples execute successfully with detailed logical reports

  4. Test CLI usage:

    pnpm build && pnpm run cli generate "Audit this video" \
      --file video.mp4 \
      --provider vertex \
      --model gemini-2.0-flash \
      --disableTools true
  5. Confirm output contains:

    • Proper frame extraction (different frames, not duplicates)
    • Logical analysis focusing on cause-and-effect
    • Clean terminal output without markdown artifacts
    • Evidence-based reporting with JSON snippets

Test Results:

  • ✅ Frame extraction works correctly with proper temporal distribution
  • ✅ Analysis returns detailed text-based logical audit in result.content
  • ✅ Both Vertex AI and Google AI Studio providers function correctly
  • ✅ Terminal output is clean (no # or ** markdown formatting)
  • ✅ Buffer and file path inputs both work correctly
  • ✅ Different prompt styles produce focused analysis

Code Quality

Have you followed code quality standards?

  • Code follows the project's style guidelines (ESLint passes)
  • Code is properly formatted (Prettier applied)
  • Self-review of code completed
  • No console.log statements (using logger instead)
  • No hardcoded API keys or secrets
  • TypeScript strict mode compliance
  • Proper error handling implemented
  • TODO/FIXME comments reference issues

Documentation

Have you updated documentation?

  • JSDoc comments added/updated for public APIs
  • README.md updated (if needed) - Not required for this feature
  • Documentation in /docs updated: docs/features/video-analysis.md completely rewritten
  • Code examples added/updated: examples/video-analysis.ts rewritten with 6 comprehensive examples
  • CHANGELOG.md updated (if applicable) - Should be updated before merge
  • Migration guide provided (if breaking changes) - Not applicable

Commit Message Format

Does your commit follow semantic commit conventions?

  • Commit message follows format: type(scope): description
  • Valid type used: feat, fix, docs, style, refactor, test, chore, build, ci, perf, revert
  • Scope specified (e.g., providers, cli, docs, middleware)

Suggested commit message:

feat(video): implement logical video analysis with Gemini 2.0 Flash

Dependencies

Does this PR add, update, or remove dependencies?

  • No dependency changes
  • Dependencies added (list below)
  • Dependencies updated (list below)
  • Dependencies removed (list below)

All required dependencies were already present:

  • fluent-ffmpeg - Already in package.json for video processing
  • @google-cloud/vertexai - Already in package.json for Vertex AI
  • @google/generative-ai - Already in package.json for Google AI Studio

External dependency:

  • FFmpeg must be installed on the system (user responsibility, documented)

Performance Impact

Does this change affect performance?

  • No performance impact
  • Performance improved (provide metrics)
  • Performance degraded (justify why acceptable)

Frame Extraction Performance:

Before: Duplicate frames bug (10 identical frames from first 0.3s)
After: Proper temporal distribution (e.g., frames at 0s, 1s, 2s, 3s...)
Improvement: Fixed critical accuracy bug - now extracts unique keyframes

Analysis Time (end-to-end):

Short videos (<10s): ~3-5 seconds
Medium videos (10-60s): ~5-10 seconds
Long videos (>60s): ~10-20 seconds

Performance primarily depends on Gemini 2.0 Flash API latency.

Frame Extraction Speed:

10s video: ~1-2 seconds to extract 10 frames
60s video: ~2-4 seconds to extract frames

Security Considerations

Are there any security implications?

  • No security implications
  • Security review needed
  • Security vulnerability fixed

Security Notes:

  • Video data is sent to Google's Gemini API (Vertex AI or Google AI Studio)
  • Users must provide their own API credentials via environment variables
  • No video data is stored, cached, or logged by NeuroLink
  • Follows existing security patterns for multimodal content
  • All video processing is done in-memory or via temporary files (auto-cleaned by FFmpeg)

Deployment Notes

Special deployment instructions?

  • Requires environment variable changes (list below)
  • No special deployment steps
  • Requires database migration
  • Requires Redis schema update
  • Other (describe below)

Required Environment Variables:

For Vertex AI:

export GOOGLE_VERTEX_PROJECT="your-project-id"
export GOOGLE_VERTEX_LOCATION="us-central1"  # optional, defaults to us-central1

For Google AI Studio:

export GOOGLE_AI_API_KEY="your-api-key"

System Requirements:

  • FFmpeg must be installed and available in system PATH
    • macOS: brew install ffmpeg
    • Linux: apt-get install ffmpeg or yum install ffmpeg
    • Windows: Download from https://ffmpeg.org/

Screenshots / Videos

If applicable, add screenshots or videos to demonstrate changes:

Example CLI Output:

🎥 NeuroLink: Critical Video Logic Audit

✅ Analyzing: Screen Recording 2026-02-13 at 12.41.38 PM.mov

Example 1: Standard Logic Audit
────────────────────────────────────────────────────────────

📋 AUDIT SUMMARY:

Strategic Overview & Intent
The video captures a screen recording of a user attempting to submit
a form. The expected logic is: User fills form → Clicks submit → 
Sees loading state → Receives confirmation or error.

The Action-Reaction Chain

Frame 0s: Initial state shows an empty form with fields...
Frame 1s: User clicks on the "Name" field...
Frame 2s: Text appears in the Name field...
Frame 3s: User clicks "Submit" button...
Frame 4s: [CRITICAL] No visual change detected...
Frame 5s: Form still visible, no feedback...

Critical Findings

1. Silent Failure Detected
{
  "frame": "4s",
  "evidence": {
    "button_state": "enabled",
    "loading_indicator": "absent",
    "error_message": "none"
  },
  "inference": "Submit action produced no visual bond"
}

Final Assessment
The workflow exhibits a critical silent failure...

Reviewer Checklist

For reviewers:

  • Code follows project style and conventions
  • Changes are well-documented
  • Tests provide adequate coverage
  • No obvious performance issues
  • No security vulnerabilities introduced
  • Breaking changes are properly documented
  • Documentation is clear and accurate

Additional Notes

Any additional information for reviewers:

Key Implementation Details:

  1. FFmpeg Integration Deep Dive:

    • The critical fix was in the temporal selection filter
    • Old (broken): select='not(mod(n\,${frameInterval}))' - extracted same frame
    • New (fixed): select='isnan(prev_selected_t)+gte(t-prev_selected_t,${intervalSec}-0.001)'
    • This ensures frames are selected based on timestamp, not frame number
  2. Provider Support:

    • Vertex AI: Uses base64 encoding for video frames
    • Google AI Studio: Uses the same base64 approach (file upload API was considered but base64 is more reliable)
    • Both use gemini-2.0-flash model (vision-capable)
  3. Prompt Engineering Philosophy:

    • The "Critical Logic Auditor" persona was designed after multiple iterations
    • Focus shifted from "what I see" to "why this happened"
    • Key prompt sections:
      • Strategic Overview (what should happen)
      • Action-Reaction Chain (frame-by-frame causality)
      • Critical Findings (evidence in JSON format)
      • Final Assessment (verdict on logical flow)
  4. Terminal Compatibility:

    • Original prompt had # Headers and **bold** which broke terminal output
    • All markdown formatting removed for clean display
    • JSON snippets are still used for evidence (properly formatted)
  5. Why Text Output Instead of Structured JSON?

    • Initial implementation used VideoAnalysisResult type with structured fields
    • User requested plain text for easier reading and copy-paste to development teams
    • Text format allows model to be more flexible and context-aware
    • Engineers can read the report directly without parsing JSON

Future Enhancement Ideas:

  • Add unit tests for frame extraction logic
  • Support for additional providers (Anthropic Claude with vision?)
  • Configurable frame extraction intervals via options
  • Video comparison mode (analyze two videos side-by-side)
  • Support for audio transcription in addition to visual analysis
  • Export analysis to PDF or HTML report format
  • Integration with bug tracking systems (auto-create issues)

Known Limitations:

  • Requires FFmpeg installed (not bundled with NeuroLink)
  • Only supports video formats that FFmpeg can process (mp4, mov, avi, webm)
  • Analysis quality depends on Gemini 2.0 Flash's vision capabilities
  • No caching of frame extraction (each run extracts frames fresh)
  • Maximum video length not enforced (but longer videos = more API cost)

Questions for Reviewers:

  1. Should we add a max video duration limit (e.g., 5 minutes)?
  2. Should frame extraction be cached to avoid re-processing?
  3. Do we need unit tests before merge, or can we add them in a follow-up PR?
  4. Should we add a --frames CLI option to override frame count?

Pre-submission Checklist

Before submitting, ensure you have:

  • Read and followed the Contributing Guidelines
  • Verified all automated pre-commit checks pass
  • Tested changes locally with pnpm test - Manual testing completed, unit tests pending
  • Built the project successfully with pnpm build
  • Run pnpm run validate:all and all checks pass - To be run before merge
  • Reviewed your own code for obvious issues
  • Ensured commit messages follow semantic format
  • Updated relevant documentation
  • Added tests for new functionality - Pending, can be follow-up PR
  • Checked that CI/CD pipeline passes (after creating PR) - Will verify after PR creation

Thank you for contributing to NeuroLink!

Summary by CodeRabbit

  • New Features

    • Introduced comprehensive video analysis capabilities supporting multiple analysis scenarios: logic audits, UI/UX bug detection, silent failure identification, workflow validation, and performance assessments.
    • Added files parameter to input options for automatic file type detection, including video files.
    • Video analysis now supports multiple AI providers for flexible deployment.
  • Documentation

    • Added detailed video analysis feature guide with usage examples for CLI and SDK.
    • Included example code demonstrating six distinct video analysis workflows.

@vercel

vercel Bot commented Feb 17, 2026

Copy link
Copy Markdown

@rishikagudla is attempting to deploy a commit to the Sachin Sharma's projects Team on Vercel.

A member of the Team first needs to authorize it.

@coderabbitai

coderabbitai Bot commented Feb 17, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Walkthrough

This PR adds end-to-end video analysis capabilities to NeuroLink, enabling automatic analysis of video files through Vertex AI and Gemini 2.0 Flash models. It includes frame extraction, intent-based feature selection, unified result formatting, core provider integration, comprehensive documentation, and usage examples.

Changes

Cohort / File(s) Summary
Documentation & Examples
docs/features/video-analysis.md, examples/video-analysis.ts
New feature documentation describing video analysis capabilities, workflow, and output structure; TypeScript example demonstrating six distinct analysis scenarios (logic audit, silent failures, UI/UX bugs, workflow validation, performance, error handling) with various provider configurations.
Implementation Design
memory-bank/video-analysis-implementation-plan.md
Comprehensive design document outlining architecture for video analysis, including intent parser module, Google Cloud Video Intelligence handler, baseProvider integration, error handling strategies, and planned public API surfaces for video analysis results.
Video Analysis Core
src/lib/adapters/video/videoAnalyzer.ts, src/lib/utils/videoAnalysisProcessor.ts
Core video analysis handler supporting Vertex AI and Gemini API providers; frame content transformation, provider selection, and configuration resolution. Utility module for detecting video frames in messages and orchestrating analysis execution with provider fallback logic.
Core Framework Integration
src/lib/core/baseProvider.ts
Integrates video analysis into generation flow; detects video frames in input messages, executes analysis when present, and merges results into generation output for both streaming and non-streaming paths.
Video Frame Processing
src/lib/processors/media/VideoProcessor.ts
Updates frame extraction logic to accept and propagate interval parameters; refactors FFmpeg select filter to use interval-driven frame selection instead of per-timestamp expressions.
Type System
src/lib/types/fileTypes.ts, src/lib/types/generateTypes.ts
Extends FileType union with video entry; adds new files field to GenerateOptions and TextGenerationOptions for auto-detection of file types including video for analysis.

Sequence Diagram(s)

sequenceDiagram
    participant Client as Client/SDK
    participant BaseProvider as BaseProvider
    participant VideoProcessor as VideoProcessor
    participant VideoAnalyzer as VideoAnalyzer
    participant Provider as AI Provider<br/>(Vertex AI/Gemini)
    
    Client->>BaseProvider: generate(messages with video file)
    BaseProvider->>VideoProcessor: extractFrames(video, interval)
    VideoProcessor->>VideoProcessor: ffmpeg frame extraction
    VideoProcessor-->>BaseProvider: frames[] with metadata
    BaseProvider->>VideoAnalyzer: analyzeVideo(frames, options)
    VideoAnalyzer->>VideoAnalyzer: buildConfig (standardized prompt)
    VideoAnalyzer->>VideoAnalyzer: transformFramesToParts (base64/text)
    VideoAnalyzer->>Provider: generateContent(message, config, model)
    Provider-->>VideoAnalyzer: analysis result text
    VideoAnalyzer-->>BaseProvider: analysis output
    BaseProvider->>BaseProvider: merge into GenerateResult.content
    BaseProvider-->>Client: GenerateResult with video analysis
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

Suggested labels

released

Suggested reviewers

  • murdore
  • YasmeenOgo

Poem

🐰 Hop along now, for videos dance,
Frame by frame, a visual trance,
Gemini and Vertex join the play,
Analyzing footage every day,
NeuroLink now sees what eyes convey! 📹✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 62.50% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically describes the main change: adding video analysis support to NeuroLink, which is the primary focus of this comprehensive PR.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Tip

Issue Planner is now in beta. Read the docs and try it out! Share your feedback on Discord.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@rishikagudla
rishikagudla force-pushed the BZ-48424-add-support-for-video-analysis-in-neurolink branch from 374a64c to 596169c Compare February 17, 2026 07:33
@rishikagudla

Copy link
Copy Markdown
Contributor Author

@CodeRabbit please review my pr

@coderabbitai

coderabbitai Bot commented Feb 17, 2026

Copy link
Copy Markdown

@rishikagudla, I'll review the PR now and provide detailed feedback on the video analysis feature implementation.

✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 12

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/lib/processors/media/VideoProcessor.ts (1)

725-765: ⚠️ Potential issue | 🟠 Major

The interval-based select filter breaks extractFrameRange for non-zero startSec.

The new select expression isnan(prev_selected_t)+gte(t-prev_selected_t,${intervalSec}-0.001) always picks the first video frame (at t≈0) and then selects every intervalSec thereafter. For extractKeyframes (full-video, starting at t=0) this is correct.

However, extractFrameRange (lines 1063-1143) targets a specific [startSec, endSec] window. With the shared runFfmpegFrameExtraction, frames will be selected starting at t=0, not at startSec. For example, with startSec=60, endSec=70, frameCount=5 (interval ≈ 2.5s), the filter selects t=0, 2.5, 5, 7.5, 10 — entirely outside the desired range — then -frames:v 5 stops output.

Consider adding an -ss startSec input seek to the ffmpeg command when startSec > 0, or incorporating timestamp bounds into the select expression:

Suggested approach: add seek and time-limited select
   private runFfmpegFrameExtraction(
     videoPath: string,
     outputDir: string,
     timestamps: number[],
     intervalSec: number,
+    startSec: number = 0,
   ): Promise<void> {
     return new Promise((resolve, reject) => {
-      const selectExpr = `isnan(prev_selected_t)+gte(t-prev_selected_t,${intervalSec}-0.001)`;
+      // When startSec > 0, offset the select to skip early frames
+      const selectExpr = startSec > 0
+        ? `gte(t,${startSec})*(isnan(prev_selected_t)+gte(t-prev_selected_t,${intervalSec}-0.001))`
+        : `isnan(prev_selected_t)+gte(t-prev_selected_t,${intervalSec}-0.001)`;

Then pass startSec from extractFrameRange.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/processors/media/VideoProcessor.ts` around lines 725 - 765, The
select filter currently picks frames relative to t=0 which breaks
extractFrameRange; update runFfmpegFrameExtraction to accept a startSec (and
optionally endSec) parameter and when startSec>0 add an input seek (-ss
startSec) to ffmpegCommand (or alternatively incorporate time bounds into
selectExpr using something like between(t, startSec, endSec) combined with your
interval logic), then ensure extractFrameRange calls runFfmpegFrameExtraction
with the startSec it computes so frames are sampled from the target window
rather than from t≈0; keep the existing selectExpr behavior for startSec=0.
🧹 Nitpick comments (4)
examples/video-analysis.ts (2)

43-166: Consider wrapping generate calls with withTimeout utility.

The coding guidelines for *.ts files specify: "Wrap async operations with withTimeout utility." None of the six neurolink.generate() calls are wrapped. Video analysis can be long-running, so timeouts would guard against indefinite hangs.

As per coding guidelines, **/*.ts: "Wrap async operations with withTimeout utility."

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@examples/video-analysis.ts` around lines 43 - 166, The six neurolink.generate
calls (e.g., in Example 1..6 using neurolink.generate with VIDEO_PATH or
fs.readFileSync) are not wrapped with the withTimeout utility; update each call
to use withTimeout(...) so async operations will time out, e.g., wrap the
Promise returned by neurolink.generate in withTimeout with an appropriate
timeout value and handle timeout errors in the existing try/catch; ensure you
replace each direct neurolink.generate invocation (including result1..result6)
with the withTimeout-wrapped call and import or reference the withTimeout helper
where it’s used.

162-166: A single failure aborts all remaining examples.

The current try/catch wraps all six examples together, so if Example 2 fails, Examples 3–6 are skipped. For a demo script showcasing independent use cases, wrapping each example in its own try/catch would be more resilient and let users see results from the examples that do succeed.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@examples/video-analysis.ts` around lines 162 - 166, The top-level try/catch
currently wrapping all six example runs causes one failure to abort the rest;
refactor by enclosing each example invocation (the individual Example 1..6 call
sites) in its own try { ... } catch (error) { console.error("\n❌ Example N
failed:", error instanceof Error ? error.message : String(error)); } block, and
remove or avoid calling process.exit(1) inside those per-example catch blocks so
subsequent examples still run; keep the existing final process.exit only for a
global fatal error if needed.
src/lib/adapters/video/videoAnalyzer.ts (1)

116-125: No timeout on ai.models.generateContent() — could hang indefinitely.

The Gemini API call has no timeout protection. If the API is unresponsive, generate() will block indefinitely. As per coding guidelines, async operations should be wrapped with a timeout utility.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/adapters/video/videoAnalyzer.ts` around lines 116 - 125, The call to
ai.models.generateContent in videoAnalyzer.ts can hang because it lacks a
timeout; wrap the ai.models.generateContent(...) invocation in the project's
timeout utility (e.g., withTimeout/runWithTimeout) so the operation is aborted
after the configured limit, pass the same model/config/contents (buildConfig(),
parts) into the timed call, and ensure you handle the timeout case by
throwing/logging a clear error or returning a safe fallback from the enclosing
function (video analysis flow) to avoid leaving the routine blocked.
src/lib/utils/videoAnalysisProcessor.ts (1)

57-73: Duplicated provider-resolution logic diverges from analyzeVideo dispatcher.

Lines 57-65 pre-resolve the provider, then pass it to analyzeVideo (videoAnalyzer.ts line 278), which has its own provider resolution. The two can diverge — e.g., here GOOGLE_AI is selected when provider === AUTO && GOOGLE_AI_API_KEY exists, but analyzeVideo routes AUTO to Vertex AI unconditionally.

Also, the project field (line 69-71) is set to undefined when options.region is truthy, but getVertexConfig() inside analyzeVideoWithVertexAI re-reads it from env vars anyway, making this field effectively ignored.

Consider removing the pre-resolution and passing options through directly, letting analyzeVideo own the provider selection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/utils/videoAnalysisProcessor.ts` around lines 57 - 73, The provider
pre-resolution block (const provider = ...) and the custom project logic before
calling analyzeVideo should be removed so analyzeVideo can own provider
selection; instead pass the original options object (or at most
options.provider) and options.model through to analyzeVideo unchanged.
Specifically, delete the provider resolution and the conditional
project/location override tied to options.region, and call
analyzeVideo(messages[0], { ...options, model: options.model ||
"gemini-2.0-flash" }) so analyzeVideo (and its
analyzeVideoWithVertexAI/getVertexConfig flow) uses a single source of truth for
provider and project resolution.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@examples/video-analysis.ts`:
- Around line 50-59: Fix the prompt string passed to neurolink.generate (the
input.text used when creating result1) to correct typographical issues: change
"whats" to "what's" and remove the double space before "it" (and optionally
clean up punctuation/capitalization for clarity) so the example prompt reads
polished and professional.
- Line 158: The header string logged by console.log("\n ERROR HANDLING REVIEW:")
is missing the emoji prefix used by other section headers; update that call
(console.log) to include a matching emoji prefix (e.g., "🐛 ERROR HANDLING
REVIEW:") so formatting is consistent with the other section headers like those
using "📋", "🕵️", "🔄", and "⚡".
- Around line 70-78: Fix the prompt typo and avoid synchronous full-file
buffering: update the prompt passed to neurolink.generate to replace "feedback
bond" with the intended phrase (e.g., "feedback" or "feedback loop") and replace
or annotate the fs.readFileSync(VIDEO_PATH) usage in the input.files for
neurolink.generate — either add a comment that synchronous Buffer loading is for
demonstration only or switch to a path-based approach (pass VIDEO_PATH as a
string or stream) to prevent loading large videos into memory; reference
neurolink.generate, VIDEO_PATH, and fs.readFileSync for the change.

In `@memory-bank/video-analysis-implementation-plan.md`:
- Around line 619-623: The example contains a duplicated call to
submitVideoAnalysis causing a redeclaration of const operation; remove the
redundant line so submitVideoAnalysis(request, config) is only called once and a
single const operation variable is declared (look for the two identical lines
referencing submitVideoAnalysis, const operation, request, and config and delete
the second occurrence).
- Around line 50-109: The document duplicates the "Architecture Overview" and
"Available analysis capabilities" sections (the block listing features like
`speech`, `ocr`, `labels`, `objects`, `shots`, `explicit`, `faces` and the
feature mapping table) and contains a typo "capabilitieVs"; remove the
duplicated block so only one canonical "Architecture Overview" and one
"Available analysis capabilities" list/table remain (keep the most complete
version under either the first occurrence or under "Core Components"), correct
the typo to "capabilities", and ensure the feature mapping table maps the
NeuroLink features (`speech`, `ocr`, `labels`, `objects`, `shots`, `explicit`,
`faces`) to Google Video Intelligence constants (`SPEECH_TRANSCRIPTION`,
`TEXT_DETECTION`, `LABEL_DETECTION`, `OBJECT_TRACKING`, `SHOT_CHANGE_DETECTION`,
`EXPLICIT_CONTENT_DETECTION`, `FACE_DETECTION`) so there's a single, consistent
place in the doc for this content.

In `@src/lib/adapters/video/videoAnalyzer.ts`:
- Around line 96-112: Extract the duplicate content-to-parts mapping used in
analyzeVideoWithVertexAI and analyzeVideoWithGeminiAPI into a single helper
(e.g., buildContentParts or convertFrameContentToParts) and replace both .map()
uses with that helper to remove duplication; in the helper, handle item types
"text" and "image" as before but explicitly detect Buffer or Uint8Array for
image payloads and convert them to base64 (fall back to validating strings and
throw a clear error if image data is neither string nor binary), ensure the
regex stripping of data URI prefixes still runs for string inputs, and preserve
the thrown Error for invalid item.type to keep existing validation.
- Around line 67-136: analyzeVideoWithVertexAI currently ignores options.project
and options.location because it always calls getVertexConfig(); update the
function to merge/override the config with the provided options by doing
something like: const config = await getVertexConfig(); const project =
options.project ?? config.project; const location = options.location ??
config.location; then use those merged values when constructing GoogleGenAI and
in the logger (keep options.model/default model logic unchanged); reference
analyzeVideoWithVertexAI and getVertexConfig to locate where to replace the
existing destructuring and ensure any logging reflects the final chosen
project/location.
- Around line 268-294: analyzeVideo currently routes AUTO to
analyzeVideoWithVertexAI unconditionally; change it to detect whether Vertex is
actually configured before calling analyzeVideoWithVertexAI and otherwise fall
back to Gemini when available. Specifically, in analyzeVideo check
AIProviderName.AUTO and call getVertexConfig() (or check the same condition used
inside getVertexConfig) to confirm Vertex credentials/config exist; if Vertex is
configured call analyzeVideoWithVertexAI(frames, options), otherwise if
process.env.GOOGLE_AI_API_KEY is present call analyzeVideoWithGeminiAPI(frames,
options), and only throw the final error if neither provider is available.

In `@src/lib/core/baseProvider.ts`:
- Around line 695-706: The current video-detection logic incorrectly triggers on
any image because hasVideoFrames checks for content parts with type === "image";
change the check to a more specific signal (e.g., require a video-specific
marker in message metadata or verify the message originated from
input.videoFiles / file-type video detection) so only video-extracted frames
trigger analysis (refer to hasVideoFrames and videoAnalysisProcessor.ts matching
logic). Also wrap the executeVideoAnalysis call with the withTimeout utility to
avoid hanging the generate() flow when Gemini is unresponsive (use withTimeout
around executeVideoAnalysis with an appropriate timeout value), and ensure
errors/timeouts are caught and handled/logged without replacing content when the
call fails.
- Around line 796-803: The current logic replaces the AI-generated content when
videoAnalysisResult is present, discarding executeGeneration output; update the
flow to either (A) add a new field videoAnalysis to the result type and payload
so enhancedResult keeps the original content and also includes videoAnalysis
(update GenerateResult in generateTypes.ts and set enhancedResult.videoAnalysis
= videoAnalysisResult), or (B) if video analysis should short-circuit
generation, move the videoAnalysisResult check before executeGeneration and
return early to avoid running the generation pipeline; modify the code around
executeGeneration, enhancedResult, and videoAnalysisResult accordingly to
implement one of these two behaviors.

In `@src/lib/types/fileTypes.ts`:
- Line 20: Remove the duplicate "video" literal from the FileType union
declaration so the union contains each file type only once; locate the FileType
type (the union of string literals that currently includes "video" twice) and
delete the redundant "video" entry to avoid confusion.

In `@src/lib/utils/videoAnalysisProcessor.ts`:
- Line 67: The current call to analyzeVideo only sends messages[0], dropping any
additional messages; update the logic in videoAnalysisProcessor so you collect
all relevant messages (e.g., filter messages for those containing
media/image/video fields or attachments) and either (a) call analyzeVideo with
the full array of filtered messages if analyzeVideo supports multiple inputs or
(b) iterate over the filtered messages and call analyzeVideo for each, then
merge/concatenate the results into videoAnalysisText. Ensure you update
references to analyzeVideo and the videoAnalysisText variable to handle an array
of inputs or aggregated results and preserve message order/context.

---

Outside diff comments:
In `@src/lib/processors/media/VideoProcessor.ts`:
- Around line 725-765: The select filter currently picks frames relative to t=0
which breaks extractFrameRange; update runFfmpegFrameExtraction to accept a
startSec (and optionally endSec) parameter and when startSec>0 add an input seek
(-ss startSec) to ffmpegCommand (or alternatively incorporate time bounds into
selectExpr using something like between(t, startSec, endSec) combined with your
interval logic), then ensure extractFrameRange calls runFfmpegFrameExtraction
with the startSec it computes so frames are sampled from the target window
rather than from t≈0; keep the existing selectExpr behavior for startSec=0.

---

Nitpick comments:
In `@examples/video-analysis.ts`:
- Around line 43-166: The six neurolink.generate calls (e.g., in Example 1..6
using neurolink.generate with VIDEO_PATH or fs.readFileSync) are not wrapped
with the withTimeout utility; update each call to use withTimeout(...) so async
operations will time out, e.g., wrap the Promise returned by neurolink.generate
in withTimeout with an appropriate timeout value and handle timeout errors in
the existing try/catch; ensure you replace each direct neurolink.generate
invocation (including result1..result6) with the withTimeout-wrapped call and
import or reference the withTimeout helper where it’s used.
- Around line 162-166: The top-level try/catch currently wrapping all six
example runs causes one failure to abort the rest; refactor by enclosing each
example invocation (the individual Example 1..6 call sites) in its own try { ...
} catch (error) { console.error("\n❌ Example N failed:", error instanceof Error
? error.message : String(error)); } block, and remove or avoid calling
process.exit(1) inside those per-example catch blocks so subsequent examples
still run; keep the existing final process.exit only for a global fatal error if
needed.

In `@src/lib/adapters/video/videoAnalyzer.ts`:
- Around line 116-125: The call to ai.models.generateContent in videoAnalyzer.ts
can hang because it lacks a timeout; wrap the ai.models.generateContent(...)
invocation in the project's timeout utility (e.g., withTimeout/runWithTimeout)
so the operation is aborted after the configured limit, pass the same
model/config/contents (buildConfig(), parts) into the timed call, and ensure you
handle the timeout case by throwing/logging a clear error or returning a safe
fallback from the enclosing function (video analysis flow) to avoid leaving the
routine blocked.

In `@src/lib/utils/videoAnalysisProcessor.ts`:
- Around line 57-73: The provider pre-resolution block (const provider = ...)
and the custom project logic before calling analyzeVideo should be removed so
analyzeVideo can own provider selection; instead pass the original options
object (or at most options.provider) and options.model through to analyzeVideo
unchanged. Specifically, delete the provider resolution and the conditional
project/location override tied to options.region, and call
analyzeVideo(messages[0], { ...options, model: options.model ||
"gemini-2.0-flash" }) so analyzeVideo (and its
analyzeVideoWithVertexAI/getVertexConfig flow) uses a single source of truth for
provider and project resolution.

Comment thread examples/video-analysis.ts
Comment thread examples/video-analysis.ts
Comment thread examples/video-analysis.ts Outdated
Comment thread memory-bank/video-analysis-implementation-plan.md Outdated
Comment thread memory-bank/video-analysis-implementation-plan.md Outdated
Comment thread src/lib/adapters/video/videoAnalyzer.ts
Comment thread src/lib/core/baseProvider.ts
Comment thread src/lib/core/baseProvider.ts Outdated
Comment thread src/lib/types/fileTypes.ts Outdated
Comment thread src/lib/utils/videoAnalysisProcessor.ts Outdated
@rishikagudla
rishikagudla force-pushed the BZ-48424-add-support-for-video-analysis-in-neurolink branch from 596169c to 6bb932c Compare February 17, 2026 09:04
@murdore
murdore force-pushed the BZ-48424-add-support-for-video-analysis-in-neurolink branch from 6bb932c to 9561315 Compare February 17, 2026 15:47
@murdore
murdore merged commit c35f8a8 into juspay:release Feb 17, 2026
8 of 9 checks passed
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 9.9.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants