Skip to content

feat(middleware): robust bad word filtering in guardrails and correct… - #177

Merged
murdore merged 1 commit into
juspay:releasefrom
amreetkhuntia:add-middlewares-to-streams
Sep 23, 2025
Merged

murdore merged 1 commit into
juspay:releasefrom
amreetkhuntia:add-middlewares-to-streams

Conversation

@amreetkhuntia

@amreetkhuntia amreetkhuntia commented Sep 18, 2025 •

Copy link
Copy Markdown
Contributor

Description

This PR implements robust bad word filtering in the guardrails middleware and corrects stream usage to match the latest API and types. The changes improve reliability and ensure all output text is properly filtered for configured bad words.

Type of Change

  • ✨ New feature (non-breaking change which adds functionality)
  • 🐛 Bug fix (non-breaking change which fixes an issue)
  • 💥 Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • 📚 Documentation update
  • 🧹 Code refactoring (no functional changes)
  • ⚡ Performance improvement
  • 🧪 Test coverage improvement
  • 🔧 Build/CI configuration change

Related Issues

  • Related to # (add issue numbers as needed)

Changes Made

  • Implements robust bad word filtering in guardrails middleware
  • Escapes regex terms to prevent false matches and errors
  • Handles both string and object stream chunks (e.g., content, textDelta)
  • Ensures all output text is properly filtered for configured bad words
  • Improves middleware reliability for streaming and generate calls
  • Updates stream usage to match latest API and types

AI Provider Impact

  • OpenAI
  • Anthropic
  • Google AI/Vertex
  • AWS Bedrock
  • Azure OpenAI
  • Hugging Face
  • Ollama
  • Mistral
  • All providers
  • No provider-specific changes

Component Impact

  • CLI
  • SDK
  • MCP Integration
  • Streaming
  • Tool Calling
  • Configuration
  • Documentation
  • Tests

Testing

  • Unit tests added/updated
  • Integration tests added/updated
  • E2E tests added/updated
  • Manual testing performed
  • All existing tests pass

Test Environment

  • OS: macOS
  • Node.js version: (fill in your version)
  • Package manager: (npm/pnpm/yarn, fill in your choice)

Performance Impact

  • No performance impact
  • Performance improvement
  • Minor performance impact (acceptable)
  • Significant performance impact (needs discussion)

Breaking Changes

None.

Screenshots/Demo

middleware-stream middleware-text

Checklist

  • My code follows the project's style guidelines
  • I have performed a self-review of my code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • Any dependent changes have been merged and published

Additional Notes

No additional notes.

Summary by CodeRabbit

  • New Features

    • Added demo middleware examples showcasing real-time and non-streaming content filtering.
  • Improvements

    • Middleware and guardrails now consistently apply during streaming across supported providers, enabling real-time bad-word filtering.
    • Enhanced guardrails to filter per-chunk stream updates with safer matching for more accurate masking.
    • Increased debug logging around middleware initialization, registration, and application to aid troubleshooting.
  • Documentation

    • New demo guides illustrate configuring custom middleware and guardrails, with examples of filtered vs. unfiltered outputs.

Comment thread src/lib/core/baseProvider.ts Outdated
// Get the base model
const baseModel = await this.getAISDKModel();

logger.info(`Retrieved base model for ${this.providerName}`, {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@amreetkhuntia Info messages are mostly logged. Can you see if all of the messages that you have made info should be debug or should be kept info only?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah moving to DEBUG


const apiKey = this.getApiKey();
const anthropicClient = createAnthropic({ apiKey });
const model = anthropicClient(this.modelName);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Wont this change cause issues?

@murdore
murdore requested a review from Copilot September 20, 2025 13:12

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR implements robust bad word filtering in the guardrails middleware and updates stream usage to match the latest API and types. The changes ensure all output text is properly filtered for configured bad words across all AI providers and improve middleware reliability.

Key changes include:

  • Updates all providers to use getAISDKModelWithMiddleware() for consistent middleware application
  • Enhances guardrails middleware with proper regex escaping and support for textDelta filtering
  • Adds comprehensive logging for middleware debugging

Reviewed Changes

Copilot reviewed 15 out of 17 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
src/lib/providers/*.ts Updated to use getAISDKModelWithMiddleware() for consistent middleware application
src/lib/middleware/builtin/guardrails.ts Enhanced bad word filtering with regex escaping and textDelta support
src/lib/middleware/registry.ts Added debug logging for middleware registration and chain building
src/lib/middleware/factory.ts Enhanced logging from debug to info level
src/lib/core/baseProvider.ts Added detailed logging for middleware application process
neurolink-demo/middleware/*.ts Added demo files showing middleware usage in text and stream scenarios

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

regex,
"*".repeat(term.length),
),
};

Copilot AI Sep 20, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code only handles textDelta chunks but doesn't filter regular string chunks anymore. The original string filtering logic was removed, which means bad words in string chunks will pass through unfiltered.

Suggested change
};
};
} else if (typeof filteredChunk === "string") {
filteredChunk = filteredChunk.replace(
regex,
"*".repeat(term.length),
);

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Well the previous one was wrong since we get always objects

@@ -53,14 +53,14 @@ export function createGuardrailsMiddleware(
filteredText = filteredText?.replace(regex, "*".repeat(term.length));

Copilot AI Sep 20, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regex is created without escaping special characters, but the escapeRegExp function is only used in the stream handler. This inconsistency could cause regex errors or incorrect matching in the generate handler.

Copilot uses AI. Check for mistakes.
@amreetkhuntia
amreetkhuntia force-pushed the add-middlewares-to-streams branch 2 times, most recently from 0fec4e2 to c93045d Compare September 22, 2025 06:46
… stream usage

- Implements robust bad word filtering in guardrails middleware
- Escapes regex terms to prevent false matches and errors
- Handles both string and object stream chunks (e.g., content, textDelta)
- Ensures all output text is properly filtered for configured bad words
- Improves middleware reliability for streaming and generate calls
- Updates stream usage to match latest API and types
@amreetkhuntia
amreetkhuntia force-pushed the add-middlewares-to-streams branch from c93045d to a075dae Compare September 22, 2025 06:57
@coderabbitai

coderabbitai Bot commented Sep 22, 2025 •

Copy link
Copy Markdown

Note

Other AI code review bot(s) detected

CodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review.

Walkthrough

Adds demo middleware examples for text and streaming with guardrails; introduces middleware-aware model retrieval across providers’ streaming paths; updates guardrails stream filtering to operate on structured textDelta chunks; and adds debug logging in base provider and middleware factory/registry.

Changes

Cohort / File(s) Summary
Demo middleware examples
neurolink-demo/middleware/middleware-text.ts, neurolink-demo/middleware/middleware-stream.ts
New demo scripts showing middleware configuration with guardrails and a custom middleware; examples for generate and stream with bad-words filtering and console logging.
Middleware-aware model retrieval in providers
src/lib/providers/anthropic.ts, src/lib/providers/anthropicBaseProvider.ts, src/lib/providers/azureOpenai.ts, src/lib/providers/googleAiStudio.ts, src/lib/providers/googleVertex.ts, src/lib/providers/litellm.ts, src/lib/providers/mistral.ts, src/lib/providers/openAI.ts, src/lib/providers/openaiCompatible.ts
Providers’ streaming paths now obtain the model via getAISDKModelWithMiddleware(options) and pass it to streamText, replacing direct model usage/client construction.
Guardrails stream filtering
src/lib/middleware/builtin/guardrails.ts
Stream filtering now processes structured chunks (e.g., textDelta.textDelta), adds escapeRegExp for safe matching, and preserves non-textDelta chunks; non-stream generate filtering unchanged.
Base provider middleware logging
src/lib/core/baseProvider.ts
Adds debug logs around base model retrieval, middleware extraction, and application within getAISDKModelWithMiddleware().
Middleware factory/registry logging
src/lib/middleware/factory.ts, src/lib/middleware/registry.ts
Adds debug logs for factory initialization, custom middleware registration, registry build steps, and chain composition.

Sequence Diagram(s)

sequenceDiagram
  autonumber
  actor App
  participant Provider as Provider.executeStream
  participant Base as BaseProvider.getAISDKModelWithMiddleware
  participant Factory as MiddlewareFactory/Registry
  participant Model as AI SDK Model
  participant Guard as Guardrails Middleware

  App->>Provider: stream(input, options)
  Provider->>Base: getAISDKModelWithMiddleware(options)
  Base->>Factory: build middleware chain (options)
  Note right of Factory: Logs registration and chain build
  Factory-->>Base: middleware-wrapped model
  Base-->>Provider: Model
  Provider->>Model: streamText(messages, tools, ...)
  Model->>Guard: onChunk(textDelta)
  rect rgba(230,245,255,0.6)
    Note over Guard: Filter bad words in textDelta<br/>(escape regex, mask terms)
  end
  Guard-->>Model: filtered chunk
  Model-->>Provider: streamed chunks
  Provider-->>App: filtered stream

  alt Contains restricted terms
    Note over App: Receives masked output
  else Clean content
    Note over App: Receives unmodified output
  end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested reviewers

  • murdore

Poem

I tap my paws on streaming trails,
Guardrails nibble naughty tales.
Middleware stacks, a burrowed thread,
Models fetched the smarter spread.
Logs like stars in midnight’s dome—
Filtered whispers hop back home.
🐇✨

Pre-merge checks and finishing touches

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 66.67% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title Check ✅ Passed The title "feat(middleware): robust bad word filtering in guardrails and correct…" accurately references the primary changes in this PR—improved guardrails bad-word filtering (regex escaping and stream/textDelta handling) and fixes to stream usage across providers—so it is relevant and specific to the changeset. It is concise and follows a conventional commit-style prefix. The trailing ellipsis suggests truncation but does not make the intent misleading.
✨ Finishing touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/lib/providers/anthropicBaseProvider.ts (1)

101-126: Stream path bypasses built‑in/MCP tools; align with other providers.

Unlike Azure/Vertex/etc., this uses options.tools directly and always sets toolChoice:"auto". That skips direct tools and MCP tools from getAllTools(), leading to inconsistent tool availability across providers.

Apply this refactor to match others:

-      const result = await streamText({
-        model,
-        prompt: options.input.text,
-        system: options.systemPrompt,
-        temperature: options.temperature,
-        maxTokens: options.maxTokens, // No default limit - unlimited unless specified
-        tools: options.tools,
-        toolChoice: "auto",
+      const shouldUseTools = !options.disableTools && this.supportsTools();
+      const baseTools = shouldUseTools ? await this.getAllTools() : {};
+      const tools = shouldUseTools
+        ? { ...baseTools, ...(options.tools || {}) }
+        : undefined;
+
+      const result = await streamText({
+        model,
+        // Prefer messages for parity with other providers/tooling
+        messages: [
+          ...(options.systemPrompt ? [{ role: "system", content: options.systemPrompt }] : []),
+          { role: "user", content: options.input.text },
+        ],
+        temperature: options.temperature,
+        maxTokens: options.maxTokens,
+        ...(tools && { tools, toolChoice: "auto", maxSteps: options.maxSteps || DEFAULT_MAX_STEPS }),
         abortSignal: timeoutController?.controller.signal,
         onStepFinish: ({ toolCalls, toolResults }) => {
           this.handleToolExecutionStorage(
             toolCalls,
             toolResults,
             options,
           ).catch((error: unknown) => {
             logger.warn(
               "[AnthropicBaseProvider] Failed to store tool executions",
               {
                 provider: this.providerName,
                 error: error instanceof Error ? error.message : String(error),
               },
             );
           });
         },
       });

Additionally, add this import at the top if you switch to DEFAULT_MAX_STEPS:

import { DEFAULT_MAX_STEPS } from "../core/constants.js";
🧹 Nitpick comments (7)
src/lib/providers/googleAiStudio.ts (1)

132-134: Redundant env var mutation; prefer avoiding side effects.

Since getAISDKModel() already uses the apiKey passed to createGoogleGenerativeAI, setting GOOGLE_GENERATIVE_AI_API_KEY here is unnecessary and can surprise embedders running multiple providers in one process. Consider removing.

-    if (!process.env.GOOGLE_GENERATIVE_AI_API_KEY) {
-      process.env.GOOGLE_GENERATIVE_AI_API_KEY = apiKey;
-    }
+    // No-op: apiKey is passed directly to the SDK; avoid mutating global env here.
src/lib/core/baseProvider.ts (2)

902-907: Trim debug payload to avoid noisy logs and accidental sensitive data.

Logging the entire middlewareOptions object can be verbose and may include configs you’d rather not print. Log a summary instead.

-    logger.debug(`Middleware extraction result`, {
-      provider: this.providerName,
-      model: this.modelName,
-      middlewareOptions,
-    });
+    logger.debug(`Middleware extraction result`, {
+      provider: this.providerName,
+      model: this.modelName,
+      enabled: middlewareOptions?.enabledMiddleware,
+      preset: middlewareOptions?.preset,
+      hasGuardrails: !!middlewareOptions?.middlewareConfig?.guardrails,
+    });

913-917: Same here: summarize instead of dumping full config.

Keep the logs readable and safe by emitting only high‑signal fields.

-      logger.debug(`Applying middleware to ${this.providerName} model`, {
-        provider: this.providerName,
-        model: this.modelName,
-        middlewareOptions,
-      });
+      logger.debug(`Applying middleware to ${this.providerName} model`, {
+        provider: this.providerName,
+        model: this.modelName,
+        enabled: middlewareOptions?.enabledMiddleware,
+        preset: middlewareOptions?.preset,
+      });
src/lib/providers/googleVertex.ts (1)

836-836: Clarify the inline comment.

The network call occurs when streamText executes, not when obtaining the model function. Consider rewording to avoid confusion.

-      const model = await this.getAISDKModelWithMiddleware(options); // This is where network connection happens!
+      const model = await this.getAISDKModelWithMiddleware(options); // Model prepared; network calls happen inside streamText.
neurolink-demo/middleware/middleware-stream.ts (2)

8-15: Custom middleware won’t run on stream with wrapGenerate only.

If you intend to demonstrate custom middleware in streaming, add the appropriate streaming hook (e.g., a stream wrapper supported by your middleware engine) in addition to wrapGenerate.


6-6: Top‑level await in a TS demo.

Ensure tsconfig targets ES2022+ modules or run via ts-node/tsx. Otherwise wrap in an async IIFE.

-const provider = await createAIProvider("vertex");
+(async () => {
+  const provider = await createAIProvider("vertex");
+  // ... rest of the demo ...
+})();
neurolink-demo/middleware/middleware-text.ts (1)

6-6: Same note on top‑level await.

Align build/runtime (ESM) or wrap in async IIFE for portability.

-const provider = await createAIProvider("vertex");
+(async () => {
+  const provider = await createAIProvider("vertex");
+  // ... rest of the demo ...
+})();
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 4f689b7 and a075dae.

⛔ Files ignored due to path filters (2)
  • neurolink-demo/middleware/proof/middleware-stream.png is excluded by !**/*.png
  • neurolink-demo/middleware/proof/middleware-text.png is excluded by !**/*.png
📒 Files selected for processing (15)
  • neurolink-demo/middleware/middleware-stream.ts (1 hunks)
  • neurolink-demo/middleware/middleware-text.ts (1 hunks)
  • src/lib/core/baseProvider.ts (1 hunks)
  • src/lib/middleware/builtin/guardrails.ts (1 hunks)
  • src/lib/middleware/factory.ts (1 hunks)
  • src/lib/middleware/registry.ts (2 hunks)
  • src/lib/providers/anthropic.ts (1 hunks)
  • src/lib/providers/anthropicBaseProvider.ts (1 hunks)
  • src/lib/providers/azureOpenai.ts (1 hunks)
  • src/lib/providers/googleAiStudio.ts (1 hunks)
  • src/lib/providers/googleVertex.ts (1 hunks)
  • src/lib/providers/litellm.ts (1 hunks)
  • src/lib/providers/mistral.ts (1 hunks)
  • src/lib/providers/openAI.ts (1 hunks)
  • src/lib/providers/openaiCompatible.ts (1 hunks)
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-09-17T17:55:15.261Z
Learnt from: RajuSudhar
PR: juspay/neurolink#173
File: src/lib/index.ts:16-16
Timestamp: 2025-09-17T17:55:15.261Z
Learning: In src/lib/types/providers.ts, ProviderConfig was renamed to AIModelProviderConfig to deduplicate type names, as there was an existing ProviderConfig type that better suited the "ProviderConfig" name. This was an intentional breaking change for better type organization.

Applied to files:

  • src/lib/core/baseProvider.ts
  • src/lib/providers/anthropicBaseProvider.ts
  • src/lib/providers/openaiCompatible.ts
  • src/lib/providers/litellm.ts
  • src/lib/providers/googleAiStudio.ts
  • src/lib/providers/azureOpenai.ts
  • src/lib/providers/googleVertex.ts
  • src/lib/providers/openAI.ts
  • src/lib/providers/anthropic.ts
  • src/lib/providers/mistral.ts
📚 Learning: 2025-09-01T06:15:59.759Z
Learnt from: amreetkhuntia
PR: juspay/neurolink#133
File: src/lib/core/types.ts:208-210
Timestamp: 2025-09-01T06:15:59.759Z
Learning: The middleware?: MiddlewareFactoryOptions field is already present in both TextGenerationOptions and StreamOptions interfaces in the neurolink codebase.

Applied to files:

  • neurolink-demo/middleware/middleware-stream.ts
  • neurolink-demo/middleware/middleware-text.ts
🧬 Code graph analysis (6)
src/lib/middleware/factory.ts (1)
src/lib/utils/logger.ts (1)
  • logger (341-380)
src/lib/middleware/registry.ts (1)
src/lib/utils/logger.ts (1)
  • logger (341-380)
src/lib/core/baseProvider.ts (1)
src/lib/utils/logger.ts (1)
  • logger (341-380)
src/lib/providers/litellm.ts (2)
examples/comprehensive-demo.js (1)
  • result (55-62)
test/runAllProvidersTests.js (2)
  • result (168-168)
  • result (446-446)
src/lib/providers/openAI.ts (1)
src/lib/neurolink.ts (1)
  • streamText (2226-2253)
src/lib/providers/mistral.ts (1)
src/lib/neurolink.ts (1)
  • streamText (2226-2253)
🔇 Additional comments (16)
src/lib/middleware/registry.ts (2)

27-27: LGTM! Appropriate debug logging for middleware registration.

The debug logging provides good visibility into which middleware is being registered, helping with troubleshooting and monitoring middleware behavior.


104-104: Good debug logging for middleware chain building.

The debug logs provide valuable visibility into the middleware configuration and evaluation process. These logs will be helpful for debugging middleware chains and understanding which middleware get applied.

Also applies to: 107-107

src/lib/middleware/factory.ts (1)

65-67: LGTM! Consistent debug logging pattern.

The debug logs follow a consistent pattern with other files and provide useful information for tracking middleware factory initialization and custom middleware registration.

src/lib/middleware/builtin/guardrails.ts (2)

49-53: Fix the inconsistent regex escaping in the generate handler.

The stream handler correctly uses escapeRegExp() to safely escape regex special characters (Line 103), but the generate handler creates unescaped regex patterns. This inconsistency could cause regex errors or incorrect matching when bad words contain special regex characters.

Apply this diff to fix the inconsistency:

      if (config.badWords?.enabled && config.badWords.list) {
        let filteredText = result.text;
        for (const term of config.badWords.list) {
-          const regex = new RegExp(term, "gi");
+          const regex = new RegExp(escapeRegExp(term), "gi");
          filteredText = filteredText?.replace(regex, "*".repeat(term.length));
        }

93-96: Helper function is correctly implemented.

The escapeRegExp function properly escapes all regex special characters to prevent false matches or regex errors. The implementation follows standard JavaScript regex escaping patterns.

src/lib/providers/mistral.ts (1)

85-87: Consistent middleware-enabled model resolution pattern.

The change from this.model to await this.getAISDKModelWithMiddleware(options) is consistent with the broader pattern across all providers in this PR. This enables centralized middleware application during model resolution.

src/lib/providers/openAI.ts (1)

367-369: Consistent middleware-enabled model resolution pattern.

The change follows the same pattern as other providers, replacing direct model usage with middleware-wrapped model resolution. This ensures consistent middleware application across all providers.

src/lib/providers/anthropic.ts (1)

163-165: Consistent middleware-enabled model resolution pattern.

The implementation follows the same pattern seen across other providers, using getAISDKModelWithMiddleware(options) to enable middleware processing before streaming. This maintains consistency across the provider ecosystem.

src/lib/providers/openaiCompatible.ts (1)

230-232: Consistent middleware-enabled model resolution pattern.

The change to use getAISDKModelWithMiddleware(options) aligns with the broader migration across all providers. This ensures middleware is applied consistently during model resolution for streaming operations.

src/lib/providers/litellm.ts (1)

179-182: Consistent middleware-enabled model resolution pattern.

The implementation follows the established pattern of using getAISDKModelWithMiddleware(options) for middleware-wrapped model resolution. This ensures LiteLLM provider has the same middleware capabilities as other providers.

src/lib/providers/googleAiStudio.ts (1)

136-136: Good switch to middleware-aware model retrieval.

Using getAISDKModelWithMiddleware(options) aligns this provider with the shared middleware pipeline and guardrails for streaming.

src/lib/core/baseProvider.ts (1)

892-898: Debug trace is useful; no functional concerns.

The added “Retrieved base model” debug context is concise and low‑risk.

src/lib/providers/azureOpenai.ts (1)

144-147: Middleware-wrapped model is correctly used for streaming.

This brings Azure’s stream path in line with the rest of the providers.

src/lib/providers/anthropicBaseProvider.ts (1)

92-92: Good: adopt middleware-aware model retrieval.

Consistent with other providers and required for guardrails on streams.

src/lib/providers/googleVertex.ts (1)

836-836: LGTM: use getAISDKModelWithMiddleware for streaming.

This ensures guardrails/middleware apply before streamText is invoked.

neurolink-demo/middleware/middleware-text.ts (1)

8-15: Demo is fine; shows middleware + guardrails on generate.

No functional issues spotted.

Comment on lines 102 to 115
for (const term of config.badWords.list) {
const regex = new RegExp(term, "gi");
if (typeof filteredChunk === "string") {
filteredChunk = filteredChunk.replace(
regex,
"*".repeat(term.length),
);
const regex = new RegExp(escapeRegExp(term), "gi");
if (
typeof filteredChunk === "object" &&
"textDelta" in filteredChunk
) {
filteredChunk = {
...filteredChunk,
textDelta: filteredChunk.textDelta.replace(
regex,
"*".repeat(term.length),
),
};
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

🧩 Analysis chain

Verify that the textDelta structure aligns with AI SDK streaming patterns.

The code processes textDelta objects which aligns with the AI SDK streaming approach. However, based on past review comments indicating the developer confirmed "we get always objects", ensure this assumption holds across all providers and streaming scenarios.


🏁 Script executed:

#!/bin/bash
# Search for textDelta usage patterns in the codebase to verify streaming object structure
echo "Searching for textDelta usage patterns..."
rg -n -C3 --type=ts "textDelta" 

echo -e "\n=== Searching for stream chunk handling patterns ==="
rg -n -C3 --type=ts "chunk.*textDelta|textDelta.*chunk"

echo -e "\n=== Checking for any remaining string-based chunk handling ==="
rg -n -C3 --type=ts "typeof.*chunk.*===.*string|string.*===.*typeof.*chunk"

Length of output: 15650


Guardrails: redact across all stream chunk shapes

Streams emit strings and objects with 'content' or 'text' as well as 'textDelta' (examples: src/lib/providers/openAI.ts, amazonBedrock.ts, sagemaker/*, src/lib/neurolink.ts). The middleware only replaces filteredChunk.textDelta (src/lib/middleware/builtin/guardrails.ts:102-115) — bad words can bypass the filter for other shapes.

  • Fix: redact in string chunks and in object chunks with 'content' or 'text' (and keep 'textDelta'); verify each field is a string before calling .replace.
  • Alternative: normalize all stream items upstream to a single canonical shape (e.g., {type: 'text-delta', textDelta: string}) and assert that invariant.

@murdore
murdore merged commit d396797 into juspay:release Sep 23, 2025
6 checks passed
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 7.42.0 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants