Skip to content

fix(pdf-processor): enforce page limits by default with actionable alternatives - #801

Merged
murdore merged 1 commit into
juspay:releasefrom
y-naaz:fix/pdf-page-limit-enforcement
Feb 3, 2026
Merged

murdore merged 1 commit into
juspay:releasefrom
y-naaz:fix/pdf-page-limit-enforcement

Conversation

@y-naaz

@y-naaz y-naaz commented Feb 2, 2026 •

Copy link
Copy Markdown

Pull Request

Description

What does this PR do?

Enforces PDF page limits by default, throwing an actionable error instead of only logging a warning. Previously, a 10,000-page PDF would be sent to the API despite exceeding limits, causing rejection, token errors, or unexpected costs.

Related Issues

Fixes #(issue number for "Page limit check only logs a warning but doesn't prevent processing")

Type of Change

Please select the type of change:

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Documentation update
  • Refactoring (no functional changes)
  • Performance improvement
  • Test coverage improvement
  • Build/CI configuration
  • Other (please describe):

Motivation and Context

Why is this change needed? What problem does it solve?

  • Problem: Lines 158-162 in pdfProcessor.ts only called logger.warn() when page limit was exceeded. Processing continued and oversized PDFs were sent to provider APIs.
  • Impact: Wasted API calls/costs, confusing provider errors, token limits exceeded without user consent
  • Solution: Add enforceLimits option (default: true) that throws an error with actionable alternatives when limits are exceeded

Changes Made

What specific changes were made?

  • Added enforceLimits?: boolean option to PDFProcessorOptions in src/lib/types/fileTypes.ts (default: true)
  • Updated page limit validation in src/lib/utils/pdfProcessor.ts to check options?.enforceLimits !== false
  • When enforcing (default): throws error with 5 specific alternatives
  • When bypassing (enforceLimits: false): logs prominent warning about bypass risks
  • Error message includes actionable alternatives:
    1. Split PDF into smaller files
    2. Extract specific pages using PDF editor
    3. Use provider with higher limits (e.g., Google AI Studio)
    4. Convert specific pages to images manually
    5. Bypass with { enforceLimits: false } (not recommended)

Breaking Changes

Does this PR introduce breaking changes?

  • No breaking changes
  • Yes, breaking changes (describe below)

This is technically a behavior change (error instead of warning), but it prevents invalid API calls that would fail anyway. Users who intentionally want to bypass can use enforceLimits: false.

Testing

How has this been tested?

  • Unit tests added/updated
  • Integration tests added/updated
  • E2E tests pass
  • Manual testing completed
  • Tested with multiple providers: [list providers]
  • Tested on multiple platforms: [list platforms]

Test Coverage

Required tests (per acceptance criteria):

  • PDF under limit succeeds
  • PDF over limit throws error
  • Error includes actionable alternatives
  • Bypass with enforceLimits: false works
  • Bypass logs prominent warning

Manual Testing Steps

  1. Process a PDF within page limits → should succeed
  2. Process a PDF exceeding page limits (default) → should throw error with alternatives
  3. Process a PDF exceeding limits with { enforceLimits: false } → should log warning and continue

Code Quality

Have you followed code quality standards?

  • Code follows the project's style guidelines (ESLint passes)
  • Code is properly formatted (Prettier applied)
  • Self-review of code completed
  • No console.log statements (using logger instead)
  • No hardcoded API keys or secrets
  • TypeScript strict mode compliance
  • Proper error handling implemented
  • TODO/FIXME comments reference issues

Documentation

Have you updated documentation?

  • JSDoc comments added/updated for public APIs
  • README.md updated (if needed)
  • Documentation in /docs updated (if needed)
  • Code examples added/updated (if needed)
  • CHANGELOG.md updated (if applicable)
  • Migration guide provided (if breaking changes)

Commit Message Format

Does your commit follow semantic commit conventions?

  • Commit message follows format: type(scope): description
  • Valid type used: feat, fix, docs, style, refactor, test, chore, build, ci, perf, revert
  • Scope specified (e.g., providers, cli, docs, middleware)

Commit: fix(pdf-processor): enforce page limits by default with actionable alternatives

Dependencies

Does this PR add, update, or remove dependencies?

  • No dependency changes
  • Dependencies added (list below)
  • Dependencies updated (list below)
  • Dependencies removed (list below)

Performance Impact

Does this change affect performance?

  • No performance impact
  • Performance improved (provide metrics)
  • Performance degraded (justify why acceptable)

Security Considerations

Are there any security implications?

  • No security implications
  • Security review needed
  • Security vulnerability fixed

This change improves cost control by preventing unintended large API calls.

Deployment Notes

Special deployment instructions?

  • No special deployment steps
  • Requires environment variable changes (list below)
  • Requires database migration
  • Requires Redis schema update
  • Other (describe below)

Screenshots / Videos

N/A - No UI changes

Reviewer Checklist

For reviewers:

  • Code follows project style and conventions
  • Changes are well-documented
  • Tests provide adequate coverage
  • No obvious performance issues
  • No security vulnerabilities introduced
  • Breaking changes are properly documented
  • Documentation is clear and accurate

Additional Notes

Any additional information for reviewers:

  • This was tagged as a "Good first issue" with simple complexity
  • Part of "Phase 1: Critical Security" milestone
  • Labels: type:bug, priority:critical, component:pdf-processor, modality:pdf

Pre-submission Checklist

Before submitting, ensure you have:

  • Read and followed the Contributing Guidelines
  • Verified all automated pre-commit checks pass
  • Tested changes locally with pnpm test
  • Built the project successfully with pnpm build
  • Run pnpm run validate:all and all checks pass
  • Reviewed your own code for obvious issues
  • Ensured commit messages follow semantic format
  • Updated relevant documentation
  • Added tests for new functionality
  • Checked that CI/CD pipeline passes (after creating PR)

Thank you for contributing to NeuroLink!

Summary by CodeRabbit

  • New Features
    • Added a configurable option to override strict PDF page limits; default behavior still enforces limits.
  • Behavior / Bug Fixes
    • When limits are enforced, processing returns a clear, specific error for page-limit exceedance; when bypassed, processing continues with a warning and guidance on risks/alternatives.
  • Tests
    • Added comprehensive tests covering enforcement, bypass behavior, and image conversion edge cases.

Copilot AI review requested due to automatic review settings February 2, 2026 13:57
@coderabbitai

coderabbitai Bot commented Feb 2, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

  • 🔍 Trigger a full review

Walkthrough

Added an optional enforceLimits flag to PDF processing that toggles between throwing a PDF page-limit error and logging a warning; also introduced a new PDF page-limit error code/factory and expanded unit tests covering enforcement and bypass behavior.

Changes

Cohort / File(s) Summary
Type Definitions
src/lib/types/fileTypes.ts
Added optional enforceLimits?: boolean to PDFProcessorOptions with JSDoc (defaults to true).
PDF Processing Logic
src/lib/utils/pdfProcessor.ts
Conditionally enforce provider page limits: when enforceLimits true, throw ErrorFactory.pdfPageLimitExceeded; when false, log a warning and continue. Added ErrorFactory import.
Error Handling
src/lib/utils/errorHandling.ts
Added PDF_PAGE_LIMIT_EXCEEDED to ERROR_CODES and a new ErrorFactory.pdfPageLimitExceeded(...) method (note: duplicate factory declaration present in the diff).
Tests
test/unit/utils/pdfProcessor.test.ts
Added tests for page-limit enforcement, bypass behavior via enforceLimits, NeuroLinkError details, and image-conversion edge cases; includes helpers and mocks.

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant PDFProcessor
  participant Provider
  participant Logger
  participant ErrorFactory

  Client->>PDFProcessor: submit PDF
  PDFProcessor->>Provider: estimatePages(pdf)
  Provider-->>PDFProcessor: estimatedPages
  alt estimatedPages > provider.maxPages and enforceLimits = true
    PDFProcessor->>ErrorFactory: pdfPageLimitExceeded(estimatedPages, maxPages, provider)
    ErrorFactory-->>PDFProcessor: NeuroLinkError
    PDFProcessor->>Client: throw NeuroLinkError
  else estimatedPages > provider.maxPages and enforceLimits = false
    PDFProcessor->>Logger: warn("bypassing page limits", details)
    PDFProcessor->>PDFProcessor: continue processing
    PDFProcessor->>Client: return result
  else within limits
    PDFProcessor->>PDFProcessor: proceed with conversion
    PDFProcessor->>Client: return result
  end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • murdore

Poem

🐰 A rabbit nibble, soft and spry,
A flag to whisper, not to cry.
Limits met, or gently passed,
Warnings hop — the checks hold fast.
Hooray for safer PDF sky! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main change: enforcing PDF page limits by default with actionable alternatives instead of just logging warnings.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@y-naaz

y-naaz commented Feb 2, 2026

Copy link
Copy Markdown
Author

@coderabbitai Trigger a Full review on this PR and check whether the changes are according to the #263

@coderabbitai

coderabbitai Bot commented Feb 2, 2026

Copy link
Copy Markdown
🧩 Analysis chain

@y-naaz, I'll fetch issue #263 to validate the changes, then trigger a full review.


🏁 Script executed:

#!/bin/bash
# Fetch issue `#263` details
gh issue view 263 --repo juspay/neurolink --json title,body,labels,state

Length of output: 2215


🏁 Script executed:

#!/bin/bash
# Fetch the PR diff to see the actual changes
gh pr diff 801 --repo juspay/neurolink

# Also get the list of files changed
echo "=== FILES CHANGED ==="
gh pr view 801 --repo juspay/neurolink --json files --jq '.files[].path'

Length of output: 2412


@y-naaz, I've reviewed the changes against issue #263. Here's my validation:

✅ Changes Match Issue Requirements

The PR implementation aligns very well with the issue specifications:

1. Core Functionality ✓

  • Added enforceLimits?: boolean to PDFProcessorOptions with proper JSDoc documentation
  • Default behavior correctly enforces limits (when enforceLimits !== false)
  • Uses the exact check pattern specified: options?.enforceLimits !== false

2. Error Enforcement ✓

The code now throws an error with exactly 5 actionable alternatives:

  1. Split PDF into smaller files
  2. Extract specific pages using PDF editor
  3. Use provider with higher limits (references Google AI Studio)
  4. Convert pages to images manually
  5. Bypass with { enforceLimits: false } (with warning)

3. Bypass Mode ✓

When enforceLimits: false, logs a prominent warning with:

  • ⚠️ emoji for visibility
  • Clear indication of risks (API rejection, token errors, costs)
  • Suggestion to consider alternatives

4. No Breaking Changes ✓

  • Backward compatible (default behavior prevents invalid API calls)
  • Only affects PDFs exceeding page limits
  • Optional bypass for edge cases

⚠️ Critical Gap: Missing Tests

The issue requires test coverage for:

  • PDF under limit succeeds
  • PDF over limit throws error
  • Error includes actionable alternatives
  • Bypass with enforceLimits: false works
  • Bypass logs prominent warning

No test files were added or modified in this PR. The PR description also confirms this with unchecked boxes in the Testing section.

Recommendation

The code implementation is excellent and matches the issue requirements, but the PR should include tests before merging to ensure:

  • The enforcement mechanism works correctly
  • Error messages are properly formatted
  • Bypass mode behaves as expected
  • No regressions in normal PDF processing

I'll now trigger a full review of the entire PR.


🧠 Learnings used
Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2026-01-31T19:48:52.730Z
Learning: Applies to src/lib/utils/pdfProcessor.ts : PDF processing must use PDFProcessor for consistent document handling across providers

Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2026-01-31T19:48:52.730Z
Learning: Applies to src/lib/utils/messageBuilder.ts : Consider token limits for multimodal content when processing images, PDFs, and CSV files
✅ Actions performed

Full review triggered.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This pull request changes PDF page limit validation from logging a warning to throwing an error by default. It adds an enforceLimits option to PDFProcessorOptions (defaulting to true) that allows users to bypass the limit with an explicit opt-in.

Changes:

  • Added enforceLimits?: boolean field to PDFProcessorOptions type with JSDoc documentation
  • Modified page limit validation logic to throw an error with actionable alternatives when limits are exceeded (default behavior)
  • Added bypass mechanism with enforceLimits: false that logs a prominent warning instead of throwing

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 2 comments.

File Description
src/lib/types/fileTypes.ts Added enforceLimits optional boolean field to PDFProcessorOptions with JSDoc explaining default behavior
src/lib/utils/pdfProcessor.ts Updated page limit validation to enforce limits by default, throwing an error with 5 actionable alternatives, or logging a warning when bypassed

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/lib/utils/pdfProcessor.ts
Comment thread src/lib/utils/pdfProcessor.ts Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@src/lib/utils/pdfProcessor.ts`:
- Around line 203-215: The new thrown Error in the PDF page-limit check should
use the SDK's ErrorFactory to create a typed error instead of raw Error; locate
the conditional that checks metadata.estimatedPages, config.maxPages and
options?.enforceLimits in pdfProcessor (the throw inside the block that
references provider and PDF_PROVIDER_CONFIGS) and replace the new Error(...)
with a call to ErrorFactory (constructing a clear error code/message and
including details like estimatedPages, maxPages, provider, and the suggested
alternatives in the error metadata or message) so the error is typed and
consistent with the rest of the SDK.
- Around line 203-222: Add unit tests for the new enforceLimits behavior in
pdfProcessor: create tests that (1) verify a PDF with metadata.estimatedPages <=
config.maxPages processes normally (under-limit), (2) assert that when
metadata.estimatedPages > config.maxPages and options.enforceLimits is not false
the function throws an Error whose message includes the detected pages,
config.maxPages and the Google AI Studio maxPages string from
PDF_PROVIDER_CONFIGS["google-ai-studio"].maxPages, and (3) assert that when
metadata.estimatedPages > config.maxPages but options.enforceLimits === false
the function does not throw and emits a logger.warn containing the bypass
warning text and page counts; use the same entry point that reads
metadata.estimatedPages/config.maxPages/options.enforceLimits and stub/mock
logger.warn to capture warnings and any network/API calls to isolate behavior.

Comment thread src/lib/utils/pdfProcessor.ts
Comment thread src/lib/utils/pdfProcessor.ts

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@src/lib/utils/pdfProcessor.ts`:
- Around line 207-215: The error message in the PDF page-limit throw is
misleading because it suggests Google AI Studio has a higher page limit; update
the thrown Error constructed in the PDF page check (the throw new Error block
that references metadata.estimatedPages, provider, config.maxPages and
PDF_PROVIDER_CONFIGS["google-ai-studio"].maxPages) to either remove the “use
Google AI Studio for higher page limits” suggestion or replace it with an
accurate note about Google AI Studio's larger file-size allowance (e.g., mention
~2000MB size limit) and keep the other alternatives and the enforceLimits hint
intact so users get correct guidance.
🧹 Nitpick comments (1)
src/lib/utils/pdfProcessor.ts (1)

206-215: Use ErrorFactory with a new method for PDF validation errors.

While the coding guidelines specify using ErrorFactory for typed errors across the SDK, there's currently no PDF-specific error method in ErrorFactory (unlike the pattern used for PPT and image validation). Consider creating a new static method like ErrorFactory.invalidPDFPageLimit(estimatedPages, maxPages, provider) to handle this case, following the same pattern as invalidPPTPages and imageTooLarge. This would both align with SDK conventions and improve error classification downstream.

Comment thread src/lib/utils/pdfProcessor.ts Outdated
@y-naaz
y-naaz force-pushed the fix/pdf-page-limit-enforcement branch 5 times, most recently from b9cfdb7 to 2326e9c Compare February 2, 2026 19:55
@y-naaz

y-naaz commented Feb 2, 2026

Copy link
Copy Markdown
Author

@coderabbitai Trigger a Full review on this PR and check whether the changes are according to the #263

@coderabbitai

coderabbitai Bot commented Feb 2, 2026

Copy link
Copy Markdown

@y-naaz, I've already validated the changes against issue #263 in my previous response. Here's a quick summary:

✅ Validation Results

The PR implementation matches issue #263 requirements very well:

  • ✅ Added enforceLimits?: boolean to PDFProcessorOptions (default: true)
  • ✅ Throws error with 5 actionable alternatives when limits are exceeded
  • ✅ Bypass mode with prominent warning when enforceLimits: false
  • ✅ No breaking changes (backward compatible)

⚠️ Missing Tests

The only gap is test coverage - no test files were added, though issue #263 requires tests for:

  • PDF under limit succeeds
  • PDF over limit throws error
  • Error includes alternatives
  • Bypass mode works correctly
  • Warning is logged in bypass mode

I'll trigger the full review again as requested.


🧠 Learnings used
Learnt from: RajuSudhar
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-11-04T22:14:18.719Z
Learning: In the juspay/neurolink repository, all new type definitions must be placed in src/lib/types/. New type definitions outside this directory should be flagged and blocked in code reviews.

Learnt from: RajuSudhar
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-11-04T22:14:18.719Z
Learning: In the juspay/neurolink repository, do not flag existing type or interface definitions located outside src/lib/types/ - these are part of a phased migration plan and will be addressed in upcoming PRs. Only enforce type centralization rules on new code going forward.

Learnt from: RajuSudhar
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2025-11-04T22:14:18.719Z
Learning: In the juspay/neurolink repository, new type definitions should use the `type` keyword instead of `interface`, unless there is a valid and justified exception. Flag new interface declarations in code reviews.
✅ Actions performed

Full review triggered.

@y-naaz

y-naaz commented Feb 2, 2026

Copy link
Copy Markdown
Author

@coderabbitai Can you check again as I have already added the test files

@coderabbitai

coderabbitai Bot commented Feb 2, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

@y-naaz, let me check the current state of the PR for test files.


🏁 Script executed:

#!/bin/bash
# Get the current list of files in the PR
echo "=== FILES IN PR ==="
gh pr view 801 --repo juspay/neurolink --json files --jq '.files[] | "\(.path) (\(.additions)+/\(.deletions)-)"'

echo -e "\n=== SEARCHING FOR TEST FILES ==="
# Search for test files in the repository
fd -e test.ts -e spec.ts -e test.js -e spec.js

# Also search in common test directories
fd . tests __tests__ test --type f 2>/dev/null || echo "No test directories found"

echo -e "\n=== SEARCHING FOR FILES WITH 'test' OR 'spec' IN NAME ==="
fd 'test|spec' --type f

Length of output: 7412


🏁 Script executed:

#!/bin/bash
# Read the test file to validate coverage
echo "=== READING TEST FILE ==="
cat test/unit/utils/pdfProcessor.test.ts

echo -e "\n\n=== READING ERROR HANDLING FILE ==="
cat src/lib/utils/errorHandling.ts

echo -e "\n\n=== GET CURRENT PR DIFF ==="
gh pr diff 801 --repo juspay/neurolink

Length of output: 56493


@y-naaz, excellent work! 🎉 You're absolutely right—the test files have been added and they provide comprehensive coverage of all requirements from issue #263.

✅ Test Coverage Validation

Your test suite (test/unit/utils/pdfProcessor.test.ts) covers all the required scenarios:

Core Requirements from Issue #263:

  1. ✅ PDF under limit succeeds - Line 43-57
  2. ✅ PDF at exactly the limit succeeds - Line 59-68
  3. ✅ PDF over limit throws error - Line 70-88
  4. ✅ Error includes actionable alternatives - Line 98-116
  5. ✅ Bypass with enforceLimits: false works - Line 131-142
  6. ✅ Bypass logs prominent warning - Line 144-164

Additional Quality Coverage:

  • ✅ Error includes correct page counts and provider info
  • ✅ NeuroLinkError with correct error code (PDF_PAGE_LIMIT_EXCEEDED)
  • ✅ Warning mentions consequences (API rejection, costs)
  • ✅ Default enforcement behavior (when not specified)
  • ✅ Explicit enforcement behavior (when set to true)
  • ✅ Different provider limits respected

✅ Supporting Infrastructure:

You also added the error handling infrastructure (src/lib/utils/errorHandling.ts):

  • PDF_PAGE_LIMIT_EXCEEDED error code
  • ErrorFactory.pdfPageLimitExceeded() method
  • Structured error with context and alternatives

🎯 Final Verdict

All requirements from issue #263 are now met:

  • ✅ Code implementation matches specification
  • ✅ Comprehensive test coverage added
  • ✅ Error handling infrastructure in place
  • ✅ No breaking changes
  • ✅ Backward compatible with bypass option

The PR is ready for merge! Great attention to detail with the test coverage. 👏


If you found this review helpful, would you consider giving us a shout-out on X?

Thank you for using CodeRabbit!

@y-naaz

y-naaz commented Feb 2, 2026

Copy link
Copy Markdown
Author

@murdore please review and merge

@y-naaz
y-naaz force-pushed the fix/pdf-page-limit-enforcement branch from 2326e9c to 90611f4 Compare February 2, 2026 20:47
…ternatives

- Add enforceLimits option to PDFProcessorOptions (default: true)

- Throw error with 5 actionable alternatives when page limit exceeded

- Allow bypass with enforceLimits: false (logs prominent warning)

Fixes: Page limit check only logged warning but did not prevent processing
@murdore
murdore force-pushed the fix/pdf-page-limit-enforcement branch from 90611f4 to 89064a8 Compare February 3, 2026 01:34
@murdore
murdore merged commit 35576aa into juspay:release Feb 3, 2026
7 checks passed
@github-actions

github-actions Bot commented Feb 3, 2026

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 9.0.1 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants