feat(multimodal): add altText support to ImageContent for accessibility - #569
Conversation
|
Important Review skippedBot user detected. To trigger a single review, invoke the You can disable this status message by setting the Note Other AI code review bot(s) detectedCodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review. Comment |
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
5d2dabc to
5a78134
Compare
✅ Single Commit Policy - COMPLIANTStatus: Policy requirements met • 1 commit • Valid format • Ready for merge 📊 View validation details📝 Commit Details
✅ Validation Results
🤖 Automated validation by NeuroLink Single Commit Enforcement |
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
There was a problem hiding this comment.
Pull request overview
This PR adds accessibility support for images by introducing an optional altText field to image content types. Alt text is automatically included as contextual information in prompts sent to AI providers, improving both accessibility (screen readers, SEO) and AI model comprehension. The implementation is backward compatible—existing code using simple Buffer or string arrays continues to work without modification.
Key Changes
- Added
ImageWithAltTexttype andaltTextfield toImageContentfor user-facing and internal representations - Updated
GenerateOptionsandStreamOptionsto accept images with alt text alongside simple image formats - Enhanced message builder to append alt text descriptions to prompts as contextual information for AI models
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 6 comments.
Show a summary per file
| File | Description |
|---|---|
src/lib/types/multimodal.ts |
Added altText field to ImageContent type and new ImageWithAltText type with comprehensive documentation |
src/lib/types/generateTypes.ts |
Updated GenerateOptions.input.images type signature to support ImageWithAltText with JSDoc examples |
src/lib/types/streamTypes.ts |
Updated StreamOptions.input.images type signature to support ImageWithAltText with JSDoc examples |
src/lib/types/content.ts |
Exported ImageWithAltText type for public API |
src/lib/utils/messageBuilder.ts |
Added helper functions to extract image data and alt text, enhanced prompt generation to include alt text as context |
src/lib/adapters/providerImageAdapter.ts |
Updated convertToContent method to handle images with alt text and preserve it in the content structure |
test/unit/multimodal/alt-text.test.ts |
Added comprehensive unit tests covering type definitions, backward compatibility, and mixed usage scenarios |
docs/features/multimodal-chat.md |
Added documentation section on alt text usage with examples and best practices |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| }); | ||
| ``` | ||
|
|
||
| !!! tip "Alt Text Best Practices" - Keep alt text concise but descriptive (under 125 characters is ideal) - Focus on the key information the image conveys - Alt text is automatically included as context in the prompt, helping AI models better understand the images |
There was a problem hiding this comment.
The tip formatting appears to be missing proper markdown list syntax. The bullets should be on separate lines with proper markdown list formatting. Consider formatting as:
!!! tip "Alt Text Best Practices"
- Keep alt text concise but descriptive (under 125 characters is ideal)
- Focus on the key information the image conveys
- Alt text is automatically included as context in the prompt, helping AI models better understand the images| !!! tip "Alt Text Best Practices" - Keep alt text concise but descriptive (under 125 characters is ideal) - Focus on the key information the image conveys - Alt text is automatically included as context in the prompt, helping AI models better understand the images | |
| !!! tip "Alt Text Best Practices" | |
| - Keep alt text concise but descriptive (under 125 characters is ideal) | |
| - Focus on the key information the image conveys | |
| - Alt text is automatically included as context in the prompt, helping AI models better understand the images |
There was a problem hiding this comment.
Updated the tip block to use proper markdown list syntax with correct 4-space indentation. The formatting now renders correctly in MkDocs.
| // Build enhanced text with alt text context for accessibility | ||
| // If any images have alt text, include them as context for the AI | ||
| const altTextDescriptions = actualImages | ||
| .map((img, idx) => | ||
| img.altText ? `[Image ${idx + 1}: ${img.altText}]` : null, | ||
| ) | ||
| .filter(Boolean); | ||
|
|
||
| const enhancedText = | ||
| altTextDescriptions.length > 0 | ||
| ? `${text}\n\nImage descriptions for context: ${altTextDescriptions.join(" ")}` | ||
| : text; |
There was a problem hiding this comment.
The alt text is being appended after the user's original text with a hardcoded label "Image descriptions for context:". This approach has a few issues:
- The appended text could confuse users if they're debugging or reviewing prompts, as it won't be clear this was added automatically
- The formatting might interfere with user's existing prompt structure or formatting
- If images fail to download/process after alt text is added, the references may become incorrect
Consider either:
- Adding the alt text descriptions at the beginning with a clearer separator (e.g.,
## Image Context\n${altTextDescriptions.join("\n")}\n\n## User Request\n${text}) - Using a more structured format that's less likely to interfere with the user's prompt
- Adding a comment in the code explaining this is intentional behavior for AI context
There was a problem hiding this comment.
Added detailed explanatory comments explaining this is intentional behavior:
- Alt text is appended because most AI providers lack native alt text fields
- Positioned at end to preserve user's original prompt structure
- Image numbering now generated BEFORE URL downloads to maintain correct references
| import { describe, it, expect } from "vitest"; | ||
| import type { | ||
| ImageContent, | ||
| ImageWithAltText, | ||
| MultimodalInput, | ||
| Content, | ||
| } from "../../../src/lib/types/multimodal.js"; | ||
|
|
||
| describe("Image Alt Text Support", () => { | ||
| describe("ImageContent type", () => { | ||
| it("should allow altText field on ImageContent", () => { | ||
| const imageContent: ImageContent = { | ||
| type: "image", | ||
| data: Buffer.from("fake-image-data"), | ||
| altText: "A product screenshot showing the main dashboard", | ||
| mediaType: "image/png", | ||
| }; | ||
|
|
||
| expect(imageContent.altText).toBe( | ||
| "A product screenshot showing the main dashboard", | ||
| ); | ||
| expect(imageContent.type).toBe("image"); | ||
| }); | ||
|
|
||
| it("should allow ImageContent without altText (backward compatible)", () => { | ||
| const imageContent: ImageContent = { | ||
| type: "image", | ||
| data: Buffer.from("fake-image-data"), | ||
| mediaType: "image/jpeg", | ||
| }; | ||
|
|
||
| expect(imageContent.altText).toBeUndefined(); | ||
| expect(imageContent.type).toBe("image"); | ||
| }); | ||
|
|
||
| it("should allow ImageContent with both altText and metadata.description", () => { | ||
| const imageContent: ImageContent = { | ||
| type: "image", | ||
| data: "data:image/png;base64,abc123", | ||
| altText: "Screen reader alt text", | ||
| mediaType: "image/png", | ||
| metadata: { | ||
| description: "Internal detailed description of the image", | ||
| quality: "high", | ||
| }, | ||
| }; | ||
|
|
||
| expect(imageContent.altText).toBe("Screen reader alt text"); | ||
| expect(imageContent.metadata?.description).toBe( | ||
| "Internal detailed description of the image", | ||
| ); | ||
| }); | ||
| }); | ||
|
|
||
| describe("ImageWithAltText type", () => { | ||
| it("should create ImageWithAltText object", () => { | ||
| const imageWithAlt: ImageWithAltText = { | ||
| data: Buffer.from("image-data"), | ||
| altText: "A chart showing quarterly revenue trends", | ||
| }; | ||
|
|
||
| expect(imageWithAlt.data).toBeInstanceOf(Buffer); | ||
| expect(imageWithAlt.altText).toBe( | ||
| "A chart showing quarterly revenue trends", | ||
| ); | ||
| }); | ||
|
|
||
| it("should allow ImageWithAltText without altText", () => { | ||
| const imageWithAlt: ImageWithAltText = { | ||
| data: "https://example.com/image.jpg", | ||
| }; | ||
|
|
||
| expect(imageWithAlt.data).toBe("https://example.com/image.jpg"); | ||
| expect(imageWithAlt.altText).toBeUndefined(); | ||
| }); | ||
|
|
||
| it("should allow string data (URL, path, or data URI)", () => { | ||
| const urlImage: ImageWithAltText = { | ||
| data: "https://example.com/image.jpg", | ||
| altText: "Remote image description", | ||
| }; | ||
|
|
||
| const pathImage: ImageWithAltText = { | ||
| data: "./images/local.png", | ||
| altText: "Local file description", | ||
| }; | ||
|
|
||
| const dataUriImage: ImageWithAltText = { | ||
| data: "data:image/png;base64,abc123", | ||
| altText: "Inline data image description", | ||
| }; | ||
|
|
||
| expect(urlImage.altText).toBe("Remote image description"); | ||
| expect(pathImage.altText).toBe("Local file description"); | ||
| expect(dataUriImage.altText).toBe("Inline data image description"); | ||
| }); | ||
| }); | ||
|
|
||
| describe("MultimodalInput with alt text", () => { | ||
| it("should accept simple images (backward compatible)", () => { | ||
| const input: MultimodalInput = { | ||
| text: "Describe this image", | ||
| images: [Buffer.from("image-data"), "https://example.com/image.jpg"], | ||
| }; | ||
|
|
||
| expect(input.images).toHaveLength(2); | ||
| }); | ||
|
|
||
| it("should accept images with alt text", () => { | ||
| const input: MultimodalInput = { | ||
| text: "Analyze these charts", | ||
| images: [ | ||
| { data: Buffer.from("chart1"), altText: "Q1 revenue chart" }, | ||
| { data: Buffer.from("chart2"), altText: "Q2 revenue chart" }, | ||
| ], | ||
| }; | ||
|
|
||
| expect(input.images).toHaveLength(2); | ||
| if (input.images) { | ||
| const img1 = input.images[0] as ImageWithAltText; | ||
| const img2 = input.images[1] as ImageWithAltText; | ||
| expect(img1.altText).toBe("Q1 revenue chart"); | ||
| expect(img2.altText).toBe("Q2 revenue chart"); | ||
| } | ||
| }); | ||
|
|
||
| it("should accept mixed images (with and without alt text)", () => { | ||
| const input: MultimodalInput = { | ||
| text: "Compare these images", | ||
| images: [ | ||
| Buffer.from("simple-image"), // Simple buffer | ||
| "https://example.com/image.jpg", // Simple URL | ||
| { data: Buffer.from("image-with-alt"), altText: "Annotated image" }, // With alt text | ||
| ], | ||
| }; | ||
|
|
||
| expect(input.images).toHaveLength(3); | ||
| }); | ||
| }); | ||
|
|
||
| describe("Content array with alt text", () => { | ||
| it("should support alt text in Content array", () => { | ||
| const content: Content[] = [ | ||
| { type: "text", text: "Analyze this image:" }, | ||
| { | ||
| type: "image", | ||
| data: Buffer.from("image-data"), | ||
| altText: "A dashboard showing KPI metrics", | ||
| mediaType: "image/png", | ||
| }, | ||
| ]; | ||
|
|
||
| expect(content).toHaveLength(2); | ||
| const imageContent = content[1] as ImageContent; | ||
| expect(imageContent.altText).toBe("A dashboard showing KPI metrics"); | ||
| }); | ||
| }); | ||
| }); |
There was a problem hiding this comment.
The test file only covers type definitions and structure, but doesn't test the actual integration behavior where alt text is appended to the prompt text in convertSimpleImagesToProviderFormat. Consider adding integration tests that verify:
- Alt text is correctly formatted and appended to the prompt (e.g.,
[Image 1: description]) - Mixed images (with and without alt text) produce the correct output format
- The enhanced text includes the "Image descriptions for context:" prefix when alt text is present
- Images without alt text don't affect the prompt text
This would ensure the actual runtime behavior matches expectations.
There was a problem hiding this comment.
Implemented 6 comprehensive integration tests covering:
- Alt text formatting with correct image numbers
- Mixed images (with/without alt text) handling
- Image numbering maintains original order
- Enhanced text includes context prefix
- Text unchanged when no alt text present
- Alt text preserved for empty buffers
| const altTextDescriptions = actualImages | ||
| .map((img, idx) => | ||
| img.altText ? `[Image ${idx + 1}: ${img.altText}]` : null, | ||
| ) | ||
| .filter(Boolean); | ||
|
|
||
| const enhancedText = | ||
| altTextDescriptions.length > 0 | ||
| ? `${text}\n\nImage descriptions for context: ${altTextDescriptions.join(" ")}` | ||
| : text; | ||
|
|
There was a problem hiding this comment.
There's a potential image numbering mismatch issue. If a URL image fails to download (line 902-912), it won't be added to actualImages, but the alt text descriptions are generated from the actualImages array indices. This means:
- If the first image is a URL that fails to download, but has alt text, it won't appear in alt text descriptions
- The numbering in
[Image ${idx + 1}: ...]might not match the order users provided their images
Consider generating the alt text descriptions before URL downloading, or tracking the original indices to maintain correct numbering even when some images fail to download.
| const altTextDescriptions = actualImages | |
| .map((img, idx) => | |
| img.altText ? `[Image ${idx + 1}: ${img.altText}]` : null, | |
| ) | |
| .filter(Boolean); | |
| const enhancedText = | |
| altTextDescriptions.length > 0 | |
| ? `${text}\n\nImage descriptions for context: ${altTextDescriptions.join(" ")}` | |
| : text; | |
| const altTextDescriptions = images | |
| .map((image, idx) => { | |
| const altText = extractAltText(image); | |
| return altText ? `[Image ${idx + 1}: ${altText}]` : null; | |
| }) | |
| .filter(Boolean); | |
| const enhancedText = | |
| altTextDescriptions.length > 0 | |
| ? `${text}\n\nImage descriptions for context: ${altTextDescriptions.join(" ")}` | |
| : text; |
There was a problem hiding this comment.
Moved alt text generation to BEFORE URL downloading. Image numbers now always match user's input order, even when some URLs fail to download.
Example: If user provides 3 images and #2 is a URL that fails:
- Old: Shows "[Image 1: ...] [Image 2: ...]" ❌ (feat: Complete Visual Ecosystem + Automated NPM Publishing v1.1.0 #2 refers to img3)
- New: Shows "[Image 1: ...] [Image 3: ...]" ✅ (matches original positions)
| images?: Array< | ||
| Buffer | string | { data: Buffer | string; altText?: string } | ||
| >, |
There was a problem hiding this comment.
[nitpick] The inline type definition { data: Buffer | string; altText?: string } duplicates the ImageWithAltText type from multimodal.ts. For better maintainability and consistency, consider importing and using ImageWithAltText:
import type { Content, ImageWithAltText } from "../types/multimodal.js";
static convertToContent(
text: string,
images?: Array<Buffer | string | ImageWithAltText>,
): Content[]This ensures type consistency across the codebase and makes future changes to the type easier to maintain.
There was a problem hiding this comment.
Imported and used ImageWithAltText type from multimodal.ts instead of inline definition. This ensures type consistency and easier maintenance.
| images?: Array< | ||
| Buffer | string | { data: Buffer | string; altText?: string } | ||
| >, |
There was a problem hiding this comment.
[nitpick] The inline type definition { data: Buffer | string; altText?: string } duplicates the ImageWithAltText type from multimodal.ts. For better maintainability and consistency, consider importing and using ImageWithAltText:
import type { Content, ImageWithAltText } from "../types/multimodal.js";
static convertToContent(
text: string,
images?: Array<Buffer | string | ImageWithAltText>,
): Content[]This ensures type consistency across the codebase and makes future changes to the type easier to maintain.
- Fix markdown formatting in documentation tip (proper list indentation) - Add explanatory comments for alt text appending behavior - Fix image numbering mismatch by generating alt text before URL downloads - Import and use ImageWithAltText type instead of inline definition - Add 6 integration tests for alt text runtime behavior Resolves PR #569 comments: - #2588463695 (markdown formatting) - #2588463738 (alt text appending explanation) - #2588463760 (integration tests) - #2588463785 (image numbering fix) - #2588463809, #2588463833 (type consistency) All tests passing: 16/16 ✓
c54e02c to
b022dd0
Compare
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
Add comprehensive alt text support for images with accessibility features: Type System Changes: - Add altText field to ImageContent type in multimodal.ts - Add ImageWithAltText helper type for user-facing API - Update GenerateOptions and StreamOptions to support ImageWithAltText - Re-export ImageWithAltText in content.ts for backward compatibility Implementation: - Update ProviderImageAdapter to import and use ImageWithAltText type - Update messageBuilder to include alt text as context in prompts - Generate alt text descriptions BEFORE URL downloads to maintain correct numbering - Add explanatory comments about alt text appending behavior Documentation: - Add "Image Alt Text for Accessibility" section to multimodal-chat.md - Include code examples showing both simple and alt-text-enabled images - Add best practices tip with proper markdown formatting - Document alt text usage in configuration and best practices sections Testing: - Add 16 comprehensive tests (10 type tests + 6 integration tests) - Test alt text formatting with correct image numbers - Test mixed images (with/without alt text) handling - Test image numbering maintains original order - Test enhanced text generation with context prefix - Test edge cases (no alt text, empty buffers) All alt-text tests passing: 16/16 ✓ Resolves: IMG-032 Fixes: #565
a683eb0 to
3b60764
Compare
🤖 AI Review & Build Compliance ✅Status: AI analysis complete • Build rules validated • Ready for review 📊 View detailed analysis results🛡️ Analysis Complete
📋 Ready for Merge When
🤖 AI analysis complete - check individual code comments for specific feedback |
|
🎉 This PR is included in version 8.6.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |

Pull Request
Description
Adds
altTextfield to image content types for accessibility (screen readers, SEO). Alt text is included as context in prompts sent to AI providers since most providers lack native alt text support.Type of Change
Related Issues
Changes Made
altTextfield toImageContenttypeImageWithAltTexttype for user-facing APIGenerateOptionsandStreamOptionsto acceptImageWithAltTextmessageBuilderto include alt text as context in promptsProviderImageAdapter.convertToContentfor alt text handlingUsage:
Backward compatible—simple
Array<Buffer | string>still works.AI Provider Impact
Component Impact
Testing
Test Environment
Performance Impact
Breaking Changes
None. Existing code continues to work unchanged.
Screenshots/Demo
N/A
Checklist
Additional Notes
Alt text is appended to the user's prompt as contextual information (e.g.,
[Image 1: Q1 revenue chart]) since AI providers don't have native alt text fields.Warning
Firewall rules blocked me from connecting to one or more addresses (expand for details)
I tried to connect to the following addresses, but was blocked by firewall rules:
googlechromelabs.github.io/usr/local/bin/node node install.mjs(dns block)https://storage.googleapis.com/chrome-for-testing-public/141.0.7390.54/linux64/chrome-headless-shell-linux64.zip/usr/local/bin/node node install.mjs(http block)If you need me to access, download, or install something from one of these locations, you can either:
Original prompt
✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.