Skip to content

fix(observability): emit generation:end exactly once on stream finalize - #989

Merged
murdore merged 1 commit into
releasefrom
fix/curator-issue-04-stream-generation-end
Apr 26, 2026
Merged

murdore merged 1 commit into
releasefrom
fix/curator-issue-04-stream-generation-end

Conversation

@murdore

@murdore murdore commented Apr 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Curator P2-4: cost listeners that subscribe to generation:end previously received zero events from sdk.stream() calls. The doc described "fires twice"; the bug on shipped 9.56.x is the opposite direction — stream() emitted stream:complete but never generation:end, leaving any listener (cost-tracker, audit log, alerting) with nothing.

Reproduction (real providers, before fix)

generate / vertex            count=1   PASS
stream  / vertex             count=0   FAIL  ← bug
generate / google-ai-studio  count=1   PASS
stream  / google-ai-studio   count=0   FAIL  ← bug
generate / litellm           count=1   PASS
stream  / litellm            count=0   FAIL  ← bug

After fix: all stream rows return count=1.

Fix

In runStandardStreamRequest's processedStream generator, emit generation:end exactly once in the finally block with the final stream state (provider, model, content, usage, finishReason, toolsUsed, prompt, temperature, maxTokens, success, error, pipelineAHandled: true). Hoist resolvedUsage to the generator scope so it's available in finally. The event payload mirrors generate()'s shape so listeners receive a consistent contract across both APIs.

Backward compatibility

Additive only. Listeners that previously received zero events on streams now receive one event with the same shape generate() produces. Any listener that already expected the event will start working; any listener that didn't will still ignore the new event.

Verification

pnpm run build
npx tsx test/continuous-test-suite-issue-04-generation-end-dedup.ts

Expected: 6 passed (3 generate + 3 stream); the OpenAI rows may SKIP/FAIL on env-specific quota / tool-injection issues unrelated to this fix.

Test plan

  • Suite passes against fixed release for vertex, google-ai-studio, litellm
  • No regression on generate() paths (still count=1)
  • pipelineAHandled: true flag preserved for Langfuse exporter dedup

Summary by CodeRabbit

Release Notes

  • Bug Fixes

    • Fixed event handling to ensure stream completion events are emitted exactly once per operation in concurrent scenarios.
  • Tests

    • Added test suite validating event emission behavior across multiple AI providers.

Copilot AI review requested due to automatic review settings April 25, 2026 10:20
@vercel

vercel Bot commented Apr 25, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
neurolink Ready Ready Preview, Comment Apr 26, 2026 7:19am

@coderabbitai

coderabbitai Bot commented Apr 25, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 6d4e9de8-6831-4763-aa38-988bc55ba20c

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Walkthrough

This PR implements per-stream concurrency-safe deduplication for generation:end emissions in streaming contexts using AsyncLocalStorage. Providers signal when they've emitted the event, preventing the orchestration layer from emitting duplicate events across concurrent streams.

Changes

Cohort / File(s) Summary
Core Stream Deduplication Logic
src/lib/neurolink.ts
Introduces AsyncLocalStorage<StreamGenerationEndContext> scope and markStreamProviderEmittedGenerationEnd() helper. Refactors stream orchestration generator to execute within this scope, tracks provider-emitted flag, and emits generation:end exactly once per stream when provider did not.
Provider Instrumentation
src/lib/providers/googleAiStudio.ts, src/lib/providers/googleVertex.ts
Each provider now imports and invokes markStreamProviderEmittedGenerationEnd() immediately before emitting generation:end, signaling the deduplication context to prevent duplicate emissions from orchestration.
Type Definitions
src/lib/types/streamDedup.ts, src/lib/types/index.ts
Introduces StreamGenerationEndContext type with providerEmitted boolean field for ALS storage, and exports it via the types barrel file.
Test Infrastructure
test/continuous-test-suite-issue-04-generation-end-dedup.ts, test/helpers/envGuard.ts
Adds continuous test script validating exactly-one generation:end emission per generate and stream call across providers, plus helper functions for conditional test skipping and provider error classification.

Sequence Diagram

sequenceDiagram
    participant Client
    participant Orchestration as Orchestration<br/>(neurolink)
    participant ALS as AsyncLocalStorage<br/>Context
    participant Provider as Native Provider<br/>(Stream)

    Client->>Orchestration: call stream()
    Orchestration->>ALS: set StreamGenerationEndContext<br/>(providerEmitted: false)
    Orchestration->>Provider: initiate stream
    Provider->>Provider: emit tokens/events
    Provider->>Provider: stream completes
    Provider->>ALS: call markStreamProviderEmittedGenerationEnd()
    ALS->>ALS: set providerEmitted: true
    Provider->>Provider: emit generation:end
    Provider-->>Client: generation:end event
    Orchestration->>ALS: check providerEmitted flag
    alt providerEmitted is false
        Orchestration->>Client: emit generation:end
    else providerEmitted is true
        Orchestration-->>Client: skip emission
    end
    Orchestration->>Client: return result
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Suggested labels

released

Suggested reviewers

  • Pdogra2520

Poem

🐰 A stream once flowed with double cheer,
Where end events would reappear,
But now with scope so neatly kept,
Each whisper emits—once, not wept! ✨

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 37.50% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically summarizes the main objective: fixing a bug where generation:end events were not being emitted for stream calls, now emitting exactly once.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/curator-issue-04-stream-generation-end

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Apr 25, 2026 •

Copy link
Copy Markdown
Contributor

✅ Single Commit Policy - COMPLIANT

Status: Policy requirements met • 1 commit • Valid format • Ready for merge

📊 View validation details

📝 Commit Details

  • Hash: 9011eddde42ff890e1b184e8297e56b8b18c4456
  • Message: fix(observability): emit generation:end exactly once on stream finalize
  • Author: Sachin Sharma

✅ Validation Results

  • Single commit requirement met
  • No merge commits in branch
  • Semantic commit message format verified
  • Ready for squash merge to release branch

🤖 Automated validation by NeuroLink Single Commit Enforcement

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@github-actions

Copy link
Copy Markdown
Contributor

Documentation Validation Results

🚀 Documentation validation passed!

Check Status Result
Frontmatter Validation ✅ Passed
TypeScript Check ✅ Passed
Build ✅ Passed
Link Validation ✅ Passed

📦 Build artifact uploaded successfully. Ready for deployment preview.

Commit: 8c6d446541a4e89421a07e49821d82784bf59c49 | Workflow: View logs

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes an observability gap where sdk.stream() never emitted the generation:end event, preventing cost/audit/alerting listeners from receiving end-of-generation data for streaming calls.

Changes:

  • Emit generation:end from the runStandardStreamRequest stream finalizer (and hoist resolvedUsage so it’s available at finalize time).
  • Add a continuous test script to count generation:end emissions for generate() vs stream() across real providers.
  • Add small test helpers and documentation capturing the Curator Issue #4 investigation and fix.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 4 comments.

File Description
src/lib/neurolink.ts Adds a generation:end emit in the stream generator finally block and hoists resolvedUsage for the final payload.
test/helpers/envGuard.ts Adds env-var skip helper and provider-credential/network error classification helper for real-provider test scripts.
test/continuous-test-suite-issue-04-generation-end-dedup.ts Adds a real-provider script that asserts exactly one generation:end emission for both generate() and stream().
docs/curator-feedback-fixes/issue-04-stream-generation-end.md Documents the reported issue, root cause, fix approach, and verification steps.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/lib/neurolink.ts Outdated
Comment on lines +6625 to +6663
// The previous implementation only fired `stream:complete`, leaving
// any subscriber to `generation:end` with zero events.
try {
const finalProvider =
metadata.fallbackProvider ?? providerName ?? "unknown";
const finalModel =
metadata.fallbackModel ??
streamModel ??
enhancedOptions.model ??
"unknown";
const finalFinishReason = streamError
? "error"
: (streamState.finishReason ?? "stop");
self.emitter.emit("generation:end", {
provider: finalProvider,
model: finalModel,
responseTime: Date.now() - streamStartTime,
toolsUsed: streamState.toolCalls?.map((t) => t.toolName),
timestamp: Date.now(),
result: {
content: accumulatedContent,
usage: resolvedUsage,
model: finalModel,
provider: finalProvider,
finishReason: finalFinishReason,
},
prompt:
enhancedOptions.input?.text ||
(enhancedOptions as Record<string, unknown>).prompt,
temperature: enhancedOptions.temperature,
maxTokens: enhancedOptions.maxTokens,
success: !streamError,
error: streamError
? streamError instanceof Error
? streamError.message
: String(streamError)
: undefined,
pipelineAHandled: true,
});
Comment thread src/lib/neurolink.ts Outdated
self.emitter.emit("generation:end", {
provider: finalProvider,
model: finalModel,
responseTime: Date.now() - streamStartTime,
Comment thread src/lib/neurolink.ts Outdated
Comment on lines +6651 to +6653
prompt:
enhancedOptions.input?.text ||
(enhancedOptions as Record<string, unknown>).prompt,
Comment on lines +41 to +49
`generation:end` exactly once with the final stream state. A
`generationEndEmitted` boolean guards against any future double-emit. The
event payload mirrors the shape used by `generate()` so listeners receive
a consistent contract.

`src/lib/neurolink.ts` lines around 6500 and 6610: hoist `resolvedUsage`
and `generationEndEmitted` to the generator's scope, populate
`resolvedUsage` after the success path resolves it, then emit in
`finally`.
@murdore
murdore force-pushed the fix/curator-issue-04-stream-generation-end branch from d1b205a to db8d093 Compare April 25, 2026 20:43
@murdore

murdore commented Apr 25, 2026

Copy link
Copy Markdown
Contributor Author

Force-pushed db8d0932 after applying the same recipe used on #988:

  • Rebased onto latest origin/release (now a66ed14c).
  • Dropped the per-issue doc docs/curator-feedback-fixes/issue-04-stream-generation-end.md from the PR (kept locally for reference; not tracked).

Final diff:

src/lib/neurolink.ts                               |  60 ++++++-
test/continuous-test-suite-issue-04-generation-end-dedup.ts | 179 +++++++++++
test/helpers/envGuard.ts                           |  43 +++++
3 files changed, 281 insertions(+), 1 deletion(-)

Suite re-verified post-rebase: 6 PASS (vertex/google-ai-studio/litellm × generate+stream), 1 SKIP (OpenAI quota), 1 FAIL pre-existing OpenAI tool-injection env issue unrelated to this fix.

@coderabbitai full review

@coderabbitai

coderabbitai Bot commented Apr 25, 2026

Copy link
Copy Markdown

@murdore, acknowledged! Triggering a full review of the rebased PR. I'll evaluate src/lib/neurolink.ts, test/continuous-test-suite-issue-04-generation-end-dedup.ts, and test/helpers/envGuard.ts against the fix objectives, the pipelineAHandled dedup contract, and the backward-compatibility guarantees described in the PR summary.


🧠 Learnings used
Learnt from: murdore
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2026-04-17T04:30:25.521Z
Learning: In neurolink (src/lib/neurolink.ts), the `pipelineAHandled: true` flag on `generation:end` events is set ONLY by `finalizeGenerateRequestResult` (the SDK-level wrapper). Native providers (amazonBedrock.ts, ollama.ts, googleAiStudio.ts, googleVertex.ts) emit their own `generation:end` events WITHOUT this flag. The `initializeMetricsListeners()` early-return guard `if (data.pipelineAHandled) return` correctly prevents duplicates: it skips the SDK-level event while allowing native provider events through to Pipeline B.
✅ Actions performed

Full review triggered.

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@murdore

murdore commented Apr 25, 2026

Copy link
Copy Markdown
Contributor Author

Force-pushed addressing reviewer Finding #1 (double-emit on native Google streams):

The native Gemini 3 paths in googleVertex.ts:2339 and googleAiStudio.ts:1042/1076 already emit generation:end themselves — they're the only Pipeline B observation source for those paths (per the comment in initializeMetricsListeners at neurolink.ts:3083: 'Native providers (Bedrock, Ollama, Gemini 3) do NOT set [pipelineAHandled] — Pipeline B remains their only observation source').

Fix: register a counter listener at the top of runStandardStreamRequest and only fire the orchestration emit when the counter is zero (no native emit happened during the stream's lifetime). This preserves:

  • AI-SDK paths: orchestration emit fires (count=1, our richer payload with pipelineAHandled: true)
  • Native paths: provider emit fires (count=1, original payload, Pipeline B observability preserved)
  • Exactly-once contract on both

Suite re-verified: 6 PASS (vertex/google-ai-studio/litellm × generate+stream count=1 each).

@coderabbitai full review

@murdore
murdore force-pushed the fix/curator-issue-04-stream-generation-end branch from db8d093 to 771e44e Compare April 25, 2026 21:13
@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@coderabbitai

coderabbitai Bot commented Apr 25, 2026

Copy link
Copy Markdown
🧩 Analysis chain

@murdore, acknowledged! Let me read the current state of the changed files before performing the full review.


🏁 Script executed:

#!/bin/bash
# Read the relevant sections of neurolink.ts around runStandardStreamRequest and initializeMetricsListeners
rg -n "runStandardStreamRequest|generation:end|pipelineAHandled|nativeEmitCount|counter|initializeMetricsListeners|processedStream|resolvedUsage|finalizeGenerateRequestResult" src/lib/neurolink.ts | head -120

echo "---FILE BOUNDARIES---"
# Get line count
wc -l src/lib/neurolink.ts

Length of output: 2202


🏁 Script executed:

#!/bin/bash
# Read the counter implementation and surrounding context
sed -n '6456,6460p' src/lib/neurolink.ts
echo "..."
sed -n '6530,6730p' src/lib/neurolink.ts

Length of output: 8339


🏁 Script executed:

#!/bin/bash
# Read the test file and envGuard helper
cat test/helpers/envGuard.ts
echo "=== TEST SUITE ==="
cat test/continuous-test-suite-issue-04-generation-end-dedup.ts

# Also check if there's a race-condition risk: does the counter listener get registered
# before the provider stream starts? Check what happens between line 6545 and the mcpStream iteration
sed -n '6460,6548p' src/lib/neurolink.ts

# Check the finalizeGenerateRequestResult to compare payload shape
sed -n '3897,3960p' src/lib/neurolink.ts

# Check second emit site around line 4046
sed -n '4040,4075p' src/lib/neurolink.ts

# Check initializeMetricsListeners guard
sed -n '3078,3120p' src/lib/neurolink.ts

Length of output: 15401


🏁 Script executed:

#!/bin/bash
# Check for potential listener leak: what createStreamResponse does with processedStream
# and whether it guarantees consumption
rg -n "createStreamResponse" src/lib/neurolink.ts | head -20

# Check if there are any early-return / non-consumption paths after processedStream is created
sed -n '6790,6830p' src/lib/neurolink.ts

# Verify the shared emitter: is `this.emitter` per-instance or truly global/shared across concurrent calls on the same instance?
rg -n "this\.emitter\s*=" src/lib/neurolink.ts | head -20

# Check concurrent stream scenario risk: does the SDK allow multiple concurrent stream() calls on the same instance?
rg -n "stream\(" src/lib/neurolink.ts | grep "public\|async stream" | head -10

# Check payload shape: does generate path include `model` at top level?
sed -n '3916,3932p' src/lib/neurolink.ts

# Check what `result` contains in the generate path to see if model is nested there
rg -n "result\.model\|textResult\.model" src/lib/neurolink.ts | head -10

Length of output: 2205


@murdore — full review of the rebased PR (post-Finding-#1 fix) below.


Overall assessment

The counter-based dedup (providerGenerationEndEmits) is a clean, well-commented solution. For the common case — a single stream per SDK instance at a time — the exactly-once contract holds correctly across AI-SDK and native-provider paths, and the pipelineAHandled interaction with initializeMetricsListeners is sound. The test suite is practical and the envGuard helper is a solid reusable addition.

One new concern surfaces from the shared-emitter architecture — details below.


Finding #2 — Concurrent-stream cross-contamination on the shared emitter

Severity: Medium (real bug on concurrent SDK usage; no impact on sequential usage)

generationEndCounter is registered on this.emitter, which is an instance-level (not call-level) emitter. Every generation:end fired by any concurrent call on the same NeuroLink instance increments the counter of every in-flight stream's counter.

Scenario:

t=0  sdk.stream({ provider: "litellm"  })   → registers counterA (AI-SDK path, expects orchestration emit)
t=1  sdk.stream({ provider: "vertex"   })   → registers counterB (native path, will self-emit)
t=5  Vertex native fires generation:end     → counterA++ AND counterB++
t=6  litellm finally: providerGenerationEndEmitsA=1 → skips orchestration emit  ← litellm event LOST
t=7  vertex  finally: providerGenerationEndEmitsB=1 → skips orchestration emit  ← correct

The inverse is equally broken when a native stream starts before an AI-SDK stream.

Minimal fix — make the counter call-scoped by checking the event's provider/model or a unique stream token:

// Option A: tag the native emit with a streamId and only count matching events
// (requires native providers to forward the tag — invasive)

// Option B: use a WeakRef / symbol stamp — overkill

// Option C: simplest — compare provider identity in the counter
const generationEndCounter = (evt: unknown): void => {
  const data = evt as Record<string, unknown>;
  // Only count events that belong to THIS stream's provider slot.
  // Native providers do not set pipelineAHandled, so any non-pipelineAHandled
  // event from our provider/model combo counts as "already emitted".
  if (!data.pipelineAHandled &&
      (data.provider === providerName || data.provider === metadata.fallbackProvider)) {
    providerGenerationEndEmits++;
  }
};

This is still imperfect if two concurrent streams share the same provider/model string (e.g. two vertex calls), but it's substantially better than the current unbounded counter and matches the real-world use pattern (unlikely to have two identical-provider concurrent streams while one is native and one is AI-SDK). A fully robust fix would require native providers to embed a per-call correlation ID in their event payload.


Finding #3 — Payload shape inconsistency: model at top level

The stream orchestration emit includes model at the top level (Line 6698); finalizeGenerateRequestResult does not (Line 3916 — model is nested inside result). Listeners that read event.model directly work for stream events but silently get undefined for generate events and must fall back to event.result.model.

This won't break existing listeners (they already handle generate events without top-level model) but is worth aligning in a follow-up so the contract is truly symmetric.


Test suite review

test/helpers/envGuard.ts — ✅ Clean, comprehensive provider-error patterns, good reusability.

test/continuous-test-suite-issue-04-generation-end-dedup.ts — ✅ Correct structure; the 250 ms grace period is sufficient for post-stream async emits.

Gap: The suite only runs targets sequentially (await between each). Finding #2 won't be caught by this suite. A supplementary concurrent test would look like:

// Run two streams in parallel on the SAME sdk instance
const sdk = new NeuroLink();
const eventsA: unknown[] = [], eventsB: unknown[] = [];
sdk.getEventEmitter().on("generation:end", (e: unknown) => {
  // distinguish by provider if payload carries it
  eventsA.push(e); eventsB.push(e);
});
await Promise.all([
  runStreamAndDrain(sdk, "litellm"),
  runStreamAndDrain(sdk, "vertex"),   // native
]);
// Assert eventsA.length === 1 AND eventsB.length === 1 (currently fails)

Non-blocking observations

  1. resolvedUsage is undefined on the error path — the generation:end emit still fires with success: false and the error string, which is correct and matches the generate() error path shape. ✅
  2. emitter.off is called before the conditional emit in finally — eliminates any chance of the counter self-incrementing on re-entrant emitters. ✅
  3. The try/catch around the emit (Line 6724) prevents a misbehaving listener from crashing the stream's cleanup. ✅

Summary: The fix is correct and a clear improvement over the pre-PR zero-emission state. Finding #2 (concurrent cross-contamination) is the only new material concern — it doesn't affect the test suite's 6 PASS because tests run sequentially. Whether to address it now or track it as a follow-up is your call given the P2-4 scope; happy to help draft the provider-identity guard if you want to land it in this PR.


🧠 Learnings used
Learnt from: murdore
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2026-04-17T04:30:25.521Z
Learning: In neurolink (src/lib/neurolink.ts), the `pipelineAHandled: true` flag on `generation:end` events is set ONLY by `finalizeGenerateRequestResult` (the SDK-level wrapper). Native providers (amazonBedrock.ts, ollama.ts, googleAiStudio.ts, googleVertex.ts) emit their own `generation:end` events WITHOUT this flag. The `initializeMetricsListeners()` early-return guard `if (data.pipelineAHandled) return` correctly prevents duplicates: it skips the SDK-level event while allowing native provider events through to Pipeline B.

Learnt from: CR
Repo: juspay/neurolink PR: 0
File: CLAUDE.md:0-0
Timestamp: 2026-04-24T03:50:12.966Z
Learning: Applies to src/lib/providers/{google,vertex}*.ts : Gemini tools and JSON schema `structuredOutput` are mutually exclusive — Google AI Studio and Vertex AI cannot use tools with `structuredOutput` simultaneously (API limitation)
✅ Actions performed

Full review triggered.

@murdore

murdore commented Apr 26, 2026

Copy link
Copy Markdown
Contributor Author

Force-pushed addressing reviewer Finding #1 follow-up (concurrency-safe dedup):

The previous fix used a counter-listener on the shared SDK emitter and would mis-attribute a concurrent unrelated stream's generation:end to this stream, suppressing the orchestration emit incorrectly.

New approach: AsyncLocalStorage scoped per-stream:

  • New streamGenerationEndContext ALS in neurolink.ts and matching StreamGenerationEndContext type in src/lib/types/streamDedup.ts (per CLAUDE.md rule 2).
  • Exported helper markStreamProviderEmittedGenerationEnd() — native providers call it from their emit sites; safe no-op outside a stream context.
  • runStandardStreamRequest wraps the processedStream IIFE in streamGenerationEndContext.run(dedupContext, ...) so any native emit within this stream's async chain flips the per-stream flag.
  • Native emit sites in googleVertex.ts:2339 and googleAiStudio.ts:1042/1076 now call markStreamProviderEmittedGenerationEnd() before emitting.
  • Concurrent streams on the same SDK instance get independent ALS scopes — they cannot suppress each other.

Suite re-verified: 6 PASS (vertex/google-ai-studio/litellm × generate+stream count=1).

@coderabbitai full review

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@coderabbitai

coderabbitai Bot commented Apr 26, 2026

Copy link
Copy Markdown
🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Read the new type file
fd "streamDedup.ts" --exec cat {}

echo "=== ALS and helper in neurolink.ts ==="
rg -n "streamGenerationEndContext|markStreamProviderEmittedGenerationEnd|StreamGenerationEndContext|streamDedup|providerEmitted|dedupContext" src/lib/neurolink.ts | head -60

echo "=== runStandardStreamRequest full region ==="
# Find line number of runStandardStreamRequest
rg -n "runStandardStreamRequest" src/lib/neurolink.ts | head -10

Length of output: 1501


🏁 Script executed:

#!/bin/bash
# Find where markStreamProviderEmittedGenerationEnd is called in googleVertex and googleAiStudio
rg -n "markStreamProviderEmittedGenerationEnd|generation:end" src/lib/providers/googleVertex.ts 2>/dev/null | head -20
rg -n "markStreamProviderEmittedGenerationEnd|generation:end" src/lib/providers/googleAiStudio.ts 2>/dev/null | head -20

# Also check if the files exist at different paths
fd "googleVertex.ts" --exec rg -n "markStreamProviderEmittedGenerationEnd|generation:end" {} | head -20
fd "googleAiStudio.ts" --exec rg -n "markStreamProviderEmittedGenerationEnd|generation:end" {} | head -20

Length of output: 2008


Good — the ALS scaffolding and call sites look structurally correct. Let me check two remaining surface areas before writing the full review.


🏁 Script executed:

#!/bin/bash
# Check amazonBedrock and ollama for stream-path generation:end emit sites
fd "amazonBedrock.ts" --exec rg -n "markStreamProviderEmittedGenerationEnd|generation:end" {} | head -20
fd "ollama.ts" --exec rg -n "markStreamProviderEmittedGenerationEnd|generation:end" {} | head -20

echo "=== ALS run() wrapping context ==="
sed -n '6560,6840p' src/lib/neurolink.ts

echo "=== markStreamProviderEmittedGenerationEnd definition ==="
sed -n '448,470p' src/lib/neurolink.ts

Length of output: 12969


@murdore — full review of the ALS-based dedup implementation below.


Overall assessment

The AsyncLocalStorage approach correctly solves Finding #2. Each stream() call gets its own isolated dedupContext; concurrent streams on the same NeuroLink instance can no longer suppress each other's orchestration emit. The mark-before-emit ordering in both Google providers, the safe no-op outside a stream context, and the factory-function wrapping pattern are all implemented correctly.

One new gap surfaces: Amazon Bedrock and Ollama stream-path native emits are not instrumented.


Finding #4 — Missing markStreamProviderEmittedGenerationEnd() on Bedrock and Ollama stream paths

Severity: Medium (double-emit on those providers' stream paths if they route through runStandardStreamRequest)

From the script output:

File Line Site Marked?
googleVertex.ts 2345 stream generation:end ✅
googleAiStudio.ts 1046 stream success generation:end ✅
googleAiStudio.ts 1082 stream failure generation:end ✅
amazonBedrock.ts 1670 stream generation:end ❌
ollama.ts 1123 stream generation:end ❌

If amazonBedrock.ts:1670 and ollama.ts:1123 are called from within the async chain wrapped by streamGenerationEndContext.run(), the dedupContext.providerEmitted flag will remain false, the finally block will fire the orchestration emit, and the listener receives two generation:end events for those providers.

Fix (mirrors what was done for Google providers):

// amazonBedrock.ts — before line 1670
markStreamProviderEmittedGenerationEnd();
streamEmitter.emit("generation:end", { … });

// ollama.ts — before line 1123
markStreamProviderEmittedGenerationEnd();
ollamaEmitter.emit("generation:end", { … });

If Bedrock/Ollama stream paths are routed outside runStandardStreamRequest and never enter the ALS scope, the mark call is a safe no-op (getStore() returns undefined) and causes no harm, so it's worth adding defensively regardless.


Correctness of the ALS implementation ✅

Propagation: streamGenerationEndContext.run(dedupContext, buildProcessedStream) invokes the factory synchronously inside the run() scope. The returned async generator inherits the ALS store through its async continuation chain — any await or yield within the generator, including those triggered by native provider code, will have access to the correct dedupContext. This is a well-established Node.js ALS behaviour.

Ordering: markStreamProviderEmittedGenerationEnd() is called before emit("generation:end", …) at every instrumented site — dedupContext.providerEmitted is therefore true before the finally block's check runs. ✅

Generate-path isolation: googleVertex.ts:2674 and googleAiStudio.ts:1335 (generate paths) do not call markStreamProviderEmittedGenerationEnd(). Since generate paths never enter the ALS scope, getStore() returns undefined and any accidental call would be a no-op anyway. ✅

Concurrent isolation: Two simultaneous sdk.stream() calls produce two independent dedupContext objects in two independent ALS scopes — they cannot observe or mutate each other's flag. Finding #2 is fully resolved. ✅


Non-blocking carry-overs

  • Finding #3 (payload shape): The orchestration emit still includes model at the top level while generate()'s emit nests it inside result{}. Still a non-breaking asymmetry; worth a follow-up normalisation.
  • The try/catch around the orchestration emit (Line ~6724) and the emitter.off before the conditional check remain sound. ✅

Bottom line: The ALS design is the right fix and is correctly implemented for the Google native paths. Please add markStreamProviderEmittedGenerationEnd() at amazonBedrock.ts:1670 and ollama.ts:1123 before merging to close the remaining double-emit gap on those providers. Happy to confirm once those two sites are instrumented.


🧠 Learnings used
Learnt from: murdore
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2026-04-17T04:30:25.521Z
Learning: In neurolink (src/lib/neurolink.ts), the `pipelineAHandled: true` flag on `generation:end` events is set ONLY by `finalizeGenerateRequestResult` (the SDK-level wrapper). Native providers (amazonBedrock.ts, ollama.ts, googleAiStudio.ts, googleVertex.ts) emit their own `generation:end` events WITHOUT this flag. The `initializeMetricsListeners()` early-return guard `if (data.pipelineAHandled) return` correctly prevents duplicates: it skips the SDK-level event while allowing native provider events through to Pipeline B.
✅ Actions performed

Full review triggered.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (2)
test/continuous-test-suite-issue-04-generation-end-dedup.ts (2)

159-166: Add a concurrent-stream scenario to actually exercise the ALS isolation.

The whole point of switching to AsyncLocalStorage (over the instance-level counter approach) was per-stream isolation across concurrent streams on the same NeuroLink instance. The current loop is strictly sequential (generate, 1s, stream, 1s, next provider), so it would also pass with the prior, non-concurrency-safe counter implementation — i.e., it cannot catch a regression of the very issue the ALS rewrite was introduced to prevent.

Consider adding a scenario that runs two sdk.stream() calls in parallel on the same sdk instance (for a provider whose native path emits generation:end, e.g. vertex or google-ai-studio) and asserts each call observes exactly one generation:end for its own correlation id (e.g. by including a unique sentinel in the prompt or by partitioning events via the per-call result/event metadata).

🧪 Sketch of a concurrent-stream check
async function countEmissionsForConcurrentStreams(
  target: ProviderTarget,
): Promise<void> {
  const testName = `stream(x2) / ${target.provider} — concurrent generation:end isolation`;
  const skip = skipIfEnvMissing(...target.envVars);
  if (skip) { record(testName, "SKIP", skip); return; }

  const sdk = new NeuroLink();
  const events: unknown[] = [];
  sdk.getEventEmitter().on("generation:end", (e) => events.push(e));
  try {
    const model = target.modelEnv ? process.env[target.modelEnv] : undefined;
    const runOne = async () => {
      const r = await sdk.stream({
        provider: target.provider as never,
        ...(model && { model }),
        input: { text: "Reply with the single word: hello" },
        maxTokens: 32,
        disableTools: true,
      } as never);
      for await (const _ of r.stream) { /* drain */ }
    };
    await Promise.all([runOne(), runOne()]);
    await new Promise((r) => setTimeout(r, 500));
    if (events.length === 2) {
      record(testName, "PASS", `expected 2, got ${events.length}`);
    } else {
      record(testName, "FAIL", `expected 2, got ${events.length}`);
    }
  } catch (err) {
    const msg = err instanceof Error ? err.message : String(err);
    isExpectedProviderError(msg)
      ? record(testName, "SKIP", msg.slice(0, 120))
      : record(testName, "FAIL", `unexpected error: ${msg.slice(0, 200)}`);
  } finally {
    await sdk.shutdown?.().catch(() => {});
  }
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-issue-04-generation-end-dedup.ts` around lines 159
- 166, The current test loop (calling countEmissionsForGenerate and
countEmissionsForStream sequentially over TARGETS) doesn't exercise
AsyncLocalStorage concurrency; add a new test function (e.g.,
countEmissionsForConcurrentStreams) that creates a single NeuroLink instance and
runs two sdk.stream() calls in parallel, listens for "generation:end" on
sdk.getEventEmitter(), drains both streams, waits briefly, and asserts you saw
exactly two generation:end events for that instance; then invoke this new
function for appropriate targets (those with native generation:end like vertex
or google-ai-studio) in main alongside the existing checks and ensure you handle
skips, provider errors, and sdk.shutdown similar to the other helpers (use
TARGETS, sdk.stream, countEmissionsForGenerate/countEmissionsForStream
references to locate insertion points).

77-157: DRY the two near-identical scenario runners.

countEmissionsForGenerate and countEmissionsForStream differ only in the sdk.generate vs sdk.stream call (and the chunk drain). A small helper that takes a runCall(sdk, target) => Promise<{ provider; model }> and a label would remove ~60 lines of duplicated SKIP/PASS/FAIL/error/shutdown plumbing and make it trivial to add the concurrent-stream variant suggested above.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-issue-04-generation-end-dedup.ts` around lines 77
- 157, The two functions countEmissionsForGenerate and countEmissionsForStream
duplicate the same SKIP/PASS/FAIL/error/shutdown/event-listening plumbing;
refactor by extracting a helper (e.g., runEmissionTest or runCall) that accepts
(target, label, runner) where runner is an async function called with the sdk
that returns { provider, model } (for stream the runner should also drain
r.stream and wait the 250ms grace period before returning). Move the shared
event subscription (sdk.getEventEmitter().on("generation:end", ...)), skip
check, try/catch error handling (isExpectedProviderError), record(...) calls and
finally sdk.shutdown?.() into the helper, and replace
countEmissionsForGenerate/countEmissionsForStream with two tiny callers that
pass a runner invoking sdk.generate(...) or sdk.stream(...) respectively.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/lib/neurolink.ts`:
- Around line 6622-6629: The stream path resolves token usage into resolvedUsage
but does not attach streamAnalytics to the synthetic "generation:end" result, so
listeners of generate() won't see stream cost data; update the code that emits
the synthetic generation:end result (the place that assigns resolvedUsage from
streamUsage/streamAnalytics) to also set result.analytics = result.analytics ??
{}; and copy the resolved streamAnalytics (or the resolved object containing
tokenUsage/cost) into result.analytics (e.g., result.analytics.cost or
result.analytics.tokenUsage as appropriate) so the emitted result mirrors the
non-stream path; apply the same change in the other stream branch around lines
where streamAnalytics is resolved (also referenced in the other block handling
stream results).
- Around line 6774-6778: The SDK span currently always sets
ATTR.GEN_AI_FINISH_REASON to "stop" for successful streams, which hides real
finish reasons; update the logic in the block that sets streamSpan.setAttribute
(around streamSpan, ATTR.GEN_AI_FINISH_REASON) to prefer
streamState.finishReason when available, otherwise fall back to setting "error"
when metadata.error or streamError is truthy, and finally "stop" as the last
fallback; ensure you still coerce to a string and handle undefined/null values.
- Around line 6828-6831: The AsyncLocalStorage context is lost because
streamGenerationEndContext.run wraps the generator creation
(buildProcessedStream) instead of the async-iterator itself; change to run the
context around the iterator so the store is active during iteration.
Specifically, call buildProcessedStream() to get the async iterator, then wrap
that iterator with a context-bound wrapper created by
streamGenerationEndContext.run(dedupContext, () => iterator) (or equivalent) so
each next()/throw()/return() invocation runs inside streamGenerationEndContext;
ensure processStreamResult still iterates the wrapped iterator and that
markStreamProviderEmittedGenerationEnd() can read/write
dedupContext.providerEmitted during iteration.

In `@test/continuous-test-suite-issue-04-generation-end-dedup.ts`:
- Around line 139-146: Replace the fragile fixed 250ms grace sleep with a
deadline-based poll: repeatedly check the events array (events.length) in a
short interval (e.g., 50ms) until events.length >= 1 or a deadline (e.g.,
2000ms) is reached, then once you observe events.length >= 1 wait an additional
stability window (~250ms) to ensure no late duplicates are emitted; after the
stability window, assert the final events.length is exactly 1 and call
record(testName, "PASS" or "FAIL", ...) using the same detail string (which
references chunks, r.provider, r.model) so the logic in this block (variables
events, chunks, r.provider, r.model, testName, record) is preserved but robust
to slow CI runners and late duplicate events.

---

Nitpick comments:
In `@test/continuous-test-suite-issue-04-generation-end-dedup.ts`:
- Around line 159-166: The current test loop (calling countEmissionsForGenerate
and countEmissionsForStream sequentially over TARGETS) doesn't exercise
AsyncLocalStorage concurrency; add a new test function (e.g.,
countEmissionsForConcurrentStreams) that creates a single NeuroLink instance and
runs two sdk.stream() calls in parallel, listens for "generation:end" on
sdk.getEventEmitter(), drains both streams, waits briefly, and asserts you saw
exactly two generation:end events for that instance; then invoke this new
function for appropriate targets (those with native generation:end like vertex
or google-ai-studio) in main alongside the existing checks and ensure you handle
skips, provider errors, and sdk.shutdown similar to the other helpers (use
TARGETS, sdk.stream, countEmissionsForGenerate/countEmissionsForStream
references to locate insertion points).
- Around line 77-157: The two functions countEmissionsForGenerate and
countEmissionsForStream duplicate the same
SKIP/PASS/FAIL/error/shutdown/event-listening plumbing; refactor by extracting a
helper (e.g., runEmissionTest or runCall) that accepts (target, label, runner)
where runner is an async function called with the sdk that returns { provider,
model } (for stream the runner should also drain r.stream and wait the 250ms
grace period before returning). Move the shared event subscription
(sdk.getEventEmitter().on("generation:end", ...)), skip check, try/catch error
handling (isExpectedProviderError), record(...) calls and finally
sdk.shutdown?.() into the helper, and replace
countEmissionsForGenerate/countEmissionsForStream with two tiny callers that
pass a runner invoking sdk.generate(...) or sdk.stream(...) respectively.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 59987b15-3a05-434f-8216-b11c3cc57684

📥 Commits

Reviewing files that changed from the base of the PR and between a66ed14 and 7e095fb.

📒 Files selected for processing (7)
  • src/lib/neurolink.ts
  • src/lib/providers/googleAiStudio.ts
  • src/lib/providers/googleVertex.ts
  • src/lib/types/index.ts
  • src/lib/types/streamDedup.ts
  • test/continuous-test-suite-issue-04-generation-end-dedup.ts
  • test/helpers/envGuard.ts

Comment thread src/lib/neurolink.ts Outdated
Comment on lines +6622 to +6629
resolvedUsage = streamUsage;
if (!resolvedUsage && streamAnalytics) {
try {
const resolved = await Promise.resolve(streamAnalytics);
if (resolved?.tokenUsage) {
resolvedUsage = resolved.tokenUsage;
}
} catch {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Carry analytics through the synthetic generation:end result.

Right now the stream path resolves usage but drops streamAnalytics from the emitted result. Any listener shared with generate() that reads data.result.analytics?.cost still won't see stream costs, so the contract is still narrower than the non-stream path.

Suggested fix
-          let resolvedUsage: unknown;
+          let resolvedUsage: unknown;
+          let resolvedAnalytics: AnalyticsData | undefined;
           try {
             for await (const chunk of mcpStream) {
               chunkCount++;
@@
-            resolvedUsage = streamUsage;
-            if (!resolvedUsage && streamAnalytics) {
+            resolvedUsage = streamUsage;
+            if (streamAnalytics) {
               try {
-                const resolved = await Promise.resolve(streamAnalytics);
-                if (resolved?.tokenUsage) {
-                  resolvedUsage = resolved.tokenUsage;
+                resolvedAnalytics = await Promise.resolve(streamAnalytics);
+                if (!resolvedUsage && resolvedAnalytics?.tokenUsage) {
+                  resolvedUsage = resolvedAnalytics.tokenUsage;
                 }
               } catch {
                 // non-blocking
               }
             }
@@
                   result: {
                     content: accumulatedContent,
                     usage: resolvedUsage,
+                    analytics: resolvedAnalytics,
                     model: finalModel,
                     provider: finalProvider,
                     finishReason: finalFinishReason,
                   },

Also applies to: 6724-6749

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 6622 - 6629, The stream path resolves
token usage into resolvedUsage but does not attach streamAnalytics to the
synthetic "generation:end" result, so listeners of generate() won't see stream
cost data; update the code that emits the synthetic generation:end result (the
place that assigns resolvedUsage from streamUsage/streamAnalytics) to also set
result.analytics = result.analytics ?? {}; and copy the resolved streamAnalytics
(or the resolved object containing tokenUsage/cost) into result.analytics (e.g.,
result.analytics.cost or result.analytics.tokenUsage as appropriate) so the
emitted result mirrors the non-stream path; apply the same change in the other
stream branch around lines where streamAnalytics is resolved (also referenced in
the other block handling stream results).

Comment thread src/lib/neurolink.ts Outdated
Comment thread src/lib/neurolink.ts Outdated
Comment on lines +6828 to +6831
const processedStream = streamGenerationEndContext.run(
dedupContext,
buildProcessedStream,
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
node <<'NODE'
const { AsyncLocalStorage } = require('async_hooks');

const als = new AsyncLocalStorage();

const makeIterator = () => (async function* () {
  console.log('first resume store =', als.getStore());
  yield 1;
  await Promise.resolve();
  console.log('second resume store =', als.getStore());
})();

(async () => {
  const it = als.run({ streamId: 'demo' }, makeIterator);
  await it.next();
  await it.next();
})();
NODE

Repository: juspay/neurolink

Length of output: 123


🏁 Script executed:

# First, let's check if the file exists and get its size
wc -l src/lib/neurolink.ts

Repository: juspay/neurolink

Length of output: 87


🏁 Script executed:

# Read the specific lines mentioned in the review comment
sed -n '6825,6835p' src/lib/neurolink.ts

Repository: juspay/neurolink

Length of output: 481


🏁 Script executed:

# Search for buildProcessedStream definition to understand what it does
rg -A 20 "const buildProcessedStream|function buildProcessedStream" src/lib/neurolink.ts

Repository: juspay/neurolink

Length of output: 962


🏁 Script executed:

# Search for streamGenerationEndContext to understand how it's created and used
rg -B 5 -A 5 "streamGenerationEndContext" src/lib/neurolink.ts

Repository: juspay/neurolink

Length of output: 1285


🏁 Script executed:

# Search for markStreamProviderEmittedGenerationEnd to understand its implementation
rg -B 5 -A 10 "markStreamProviderEmittedGenerationEnd" src/lib/neurolink.ts

Repository: juspay/neurolink

Length of output: 1248


🏁 Script executed:

# Find all call sites of markStreamProviderEmittedGenerationEnd
rg -B 3 -A 3 "markStreamProviderEmittedGenerationEnd()" src/lib/

Repository: juspay/neurolink

Length of output: 3993


🏁 Script executed:

# Also check if it's exported and used in external files
rg "markStreamProviderEmittedGenerationEnd" --type ts --type tsx

Repository: juspay/neurolink

Length of output: 88


🏁 Script executed:

# Find processStreamResult to see how it iterates the stream
rg -B 5 -A 20 "processStreamResult" src/lib/neurolink.ts | head -80

Repository: juspay/neurolink

Length of output: 1649


🏁 Script executed:

# Let me also check where dedupContext is used after the stream is processed
rg -A 10 "providerEmitted" src/lib/neurolink.ts

Repository: juspay/neurolink

Length of output: 1270


AsyncLocalStorage context is lost during stream iteration; dedup will silently fail.

The reproducer proves that AsyncLocalStorage.run() does not preserve context across async generator resumptions—both logs returned undefined instead of the store value. Since buildProcessedStream() returns an async generator that executes during iteration (inside processStreamResult()), calls to markStreamProviderEmittedGenerationEnd() from provider code will execute outside the run() context. This means dedupContext.providerEmitted will never be set, and the dedup check will always emit a duplicate generation:end event even when the provider already emitted one.

Replace streamGenerationEndContext.run(dedupContext, buildProcessedStream) with an iterator-bound context that wraps the result of buildProcessedStream() rather than its creation.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/neurolink.ts` around lines 6828 - 6831, The AsyncLocalStorage context
is lost because streamGenerationEndContext.run wraps the generator creation
(buildProcessedStream) instead of the async-iterator itself; change to run the
context around the iterator so the store is active during iteration.
Specifically, call buildProcessedStream() to get the async iterator, then wrap
that iterator with a context-bound wrapper created by
streamGenerationEndContext.run(dedupContext, () => iterator) (or equivalent) so
each next()/throw()/return() invocation runs inside streamGenerationEndContext;
ensure processStreamResult still iterates the wrapped iterator and that
markStreamProviderEmittedGenerationEnd() can read/write
dedupContext.providerEmitted during iteration.

Comment on lines +139 to +146
// small grace period for any post-stream async emit
await new Promise((r) => setTimeout(r, 250));
const detail = `count=${events.length}; chunks=${chunks}; provider=${r.provider}; model=${r.model}`;
if (events.length === 1) {
record(testName, "PASS", `expected 1, ${detail}`);
} else {
record(testName, "FAIL", `expected 1, got ${events.length}: ${detail}`);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Fixed 250 ms grace period is fragile.

A hard-coded setTimeout(250) after stream drain may be too short on slow CI runners or for providers that flush generation:end from a microtask chain after the last chunk, producing a flaky count=0 FAIL. Conversely, if a buggy provider emits a duplicate >250 ms later, the test will pass and miss it.

Consider polling for the expected count with a deadline (e.g., wait until events.length >= 1 or a 2 s timeout, then verify it stays at 1 for an additional ~250 ms window) instead of a single fixed sleep.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@test/continuous-test-suite-issue-04-generation-end-dedup.ts` around lines 139
- 146, Replace the fragile fixed 250ms grace sleep with a deadline-based poll:
repeatedly check the events array (events.length) in a short interval (e.g.,
50ms) until events.length >= 1 or a deadline (e.g., 2000ms) is reached, then
once you observe events.length >= 1 wait an additional stability window (~250ms)
to ensure no late duplicates are emitted; after the stability window, assert the
final events.length is exactly 1 and call record(testName, "PASS" or "FAIL",
...) using the same detail string (which references chunks, r.provider, r.model)
so the logic in this block (variables events, chunks, r.provider, r.model,
testName, record) is preserved but robust to slow CI runners and late duplicate
events.

@murdore

murdore commented Apr 26, 2026

Copy link
Copy Markdown
Contributor Author

Force-pushed addressing the reviewer's concern that AsyncLocalStorage doesn't reliably propagate through async-generator yield boundaries.

New approach: drop ALS entirely. Attach a fresh mutable dedupContext object directly onto enhancedOptions._streamDedupContext. Native providers receive these options through their normal call chain and call markStreamProviderEmittedGenerationEnd(options), which mutates the same object the orchestration's finally block reads via closure capture.

Why it's concurrency-safe: each sdk.stream() call creates its own enhancedOptions object → its own dedupContext instance. Concurrent streams literally pass through different option objects so their contexts cannot interfere. No ALS, no shared listeners, no async-chain assumptions.

Wire-up:

  • src/lib/neurolink.ts:6572-6586 attaches dedupContext to enhancedOptions._streamDedupContext and reads it directly in the finally block (closure capture).
  • markStreamProviderEmittedGenerationEnd(options) now takes options and mutates options._streamDedupContext.providerEmitted = true.
  • googleVertex.ts:2345 and googleAiStudio.ts:1047/1083 pass their options / params.options to the helper.

Suite re-verified: 6 PASS (vertex/google-ai-studio/litellm × generate+stream count=1).

@coderabbitai full review

@murdore
murdore force-pushed the fix/curator-issue-04-stream-generation-end branch from 7e095fb to 9c67e68 Compare April 26, 2026 06:24
@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@coderabbitai

coderabbitai Bot commented Apr 26, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
echo "=== StreamDedup type file ==="
fd "streamDedup.ts" --exec cat {}

echo "=== markStreamProviderEmittedGenerationEnd new signature ==="
rg -n "markStreamProviderEmittedGenerationEnd\|_streamDedupContext\|dedupContext\|StreamDedup\|streamDedup" src/lib/neurolink.ts | head -40

echo "=== neurolink.ts 6560-6600 (attach dedupContext) ==="
sed -n '6560,6610p' src/lib/neurolink.ts

echo "=== neurolink.ts finally block region (look for dedupContext.providerEmitted check) ==="
rg -n "providerEmitted\|dedupContext\|pipelineAHandled" src/lib/neurolink.ts | head -30

echo "=== googleVertex.ts around 2345 ==="
sed -n '2335,2360p' src/lib/providers/googleVertex.ts

echo "=== googleAiStudio.ts around 1047 and 1083 ==="
sed -n '1038,1060p' src/lib/providers/googleAiStudio.ts
echo "..."
sed -n '1075,1095p' src/lib/providers/googleAiStudio.ts

echo "=== amazonBedrock.ts around 1670 ==="
sed -n '1660,1685p' src/lib/providers/amazonBedrock.ts

echo "=== ollama.ts around 1123 ==="
sed -n '1113,1135p' src/lib/providers/ollama.ts

echo "=== types/streamDedup.ts or _streamDedupContext type definition ==="
fd "streamDedup.ts" --exec cat {}
rg -rn "_streamDedupContext" src/lib/ | head -20

Length of output: 9887


🏁 Script executed:

#!/bin/bash
echo "=== finally block in runStandardStreamRequest ==="
# Find the finally block that emits generation:end
rg -n "providerEmitted\|dedupContext\|\.n\?" src/lib/neurolink.ts | head -30

# Read lines around the finally block
rg -n "finally" src/lib/neurolink.ts | tail -20

echo "=== Read the full finally block ==="
sed -n '6680,6770p' src/lib/neurolink.ts

echo "=== Does Bedrock route through runStandardStreamRequest? ==="
rg -n "runStandardStreamRequest\|amazonBedrock\|bedrock" src/lib/neurolink.ts | grep -i "bedrock\|runStandard" | head -20

echo "=== Does Ollama route through runStandardStreamRequest? ==="
rg -n "runStandardStreamRequest\|ollama" src/lib/neurolink.ts | grep -i "ollama\|runStandard" | head -20

echo "=== How are options passed to Bedrock provider? ==="
sed -n '1650,1695p' src/lib/providers/amazonBedrock.ts

echo "=== How are options passed to Ollama provider? ==="
sed -n '1108,1140p' src/lib/providers/ollama.ts

echo "=== stale comment in streamDedup.ts ==="
cat src/lib/types/streamDedup.ts

echo "=== markStreamProviderEmittedGenerationEnd full definition ==="
rg -n "export.*markStreamProviderEmittedGenerationEnd\|function markStreamProviderEmittedGenerationEnd\|STREAM_DEDUP" src/lib/neurolink.ts
sed -n '448,480p' src/lib/neurolink.ts

Length of output: 9800


@murdore — full review of the options-based dedup implementation below.


Overall assessment

The closure-capture approach is correct and directly sidesteps the ALS yield-boundary propagation concern. Per-call isolation is guaranteed by the fact that each sdk.stream() invocation creates its own enhancedOptions object; concurrent streams on the same NeuroLink instance carry independent _streamDedupContext references and cannot cross-contaminate. The Google provider instrumentation is properly ordered (mark before emit). No new material correctness issues — two items below warrant a quick cleanup before merging.


Item 1 — Stale "AsyncLocalStorage" comments in two places

The previous ALS-based implementation left behind stale references in two spots:

src/lib/types/streamDedup.ts — the entire JSDoc block ends with:

"AsyncLocalStorage scopes each stream's flag to its own async chain."

This contradicts the current approach. Suggested replacement:

/**
 * Curator P2-4 dedup (concurrency-safe): per-stream context that lets
 * the orchestration's `runStandardStreamRequest` finally block know
 * whether a native provider path already emitted `generation:end` for
 * THIS stream. Attached directly to `enhancedOptions._streamDedupContext`
 * so each `sdk.stream()` call carries its own isolated instance —
 * concurrent streams on the same NeuroLink instance cannot interfere.
 */
export type StreamGenerationEndContext = { providerEmitted: boolean };

src/lib/neurolink.ts finally block (around Line 6750) — comment reads:

"Per-stream flag is concurrency-safe because it's scoped via AsyncLocalStorage."

Should be updated to:

"Per-stream flag is concurrency-safe: it lives on the closed-over enhancedOptions object created uniquely per sdk.stream() call."


Item 2 — Bedrock and Ollama: confirming no action needed

From the script output, runStandardStreamRequest has zero callers that pass through the Bedrock or Ollama stack. Those providers expose independent stream paths that never enter the runStandardStreamRequest generator, so:

  • No dedupContext is created for those streams
  • The orchestration finally block never executes for those streams
  • Each provider emits generation:end exactly once from its own code

Finding #4 from the previous iteration does not apply to the options-based approach. The markStreamProviderEmittedGenerationEnd() helper would be a harmless no-op if called there (missing _streamDedupContext field → ctx is undefined → early return), but it's not needed.


Correctness checklist

Check Status
Per-call isolation via separate enhancedOptions object ✅
Closure capture: finally reads the same dedupContext reference ✅
Mark-before-emit ordering at googleVertex.ts and googleAiStudio.ts ✅
markStreamProviderEmittedGenerationEnd safe no-op when _streamDedupContext absent ✅
pipelineAHandled: true set on orchestration emit (Pipeline A contract) ✅
try/catch around orchestration emit prevents listener crash from halting cleanup ✅
Generate path never enters the dedup path ✅
Bedrock / Ollama: bypass runStandardStreamRequest, no double-emit risk ✅

Carry-over non-blocking item

Finding #3 (payload shape asymmetry): orchestration emit still has model at the top level while generate()'s emit nests it inside result. Still non-breaking; tracking as a follow-up.


Summary: The implementation is correct and concurrency-safe. Fix the two stale "AsyncLocalStorage" comments and this is ready to merge. Happy to confirm once those are updated.


✏️ Learnings added
Learnt from: murdore
URL: https://github.com/juspay/neurolink/pull/989

Timestamp: 2026-04-26T06:27:32.928Z
Learning: In juspay/neurolink PR `#989`, the `generation:end` dedup for `sdk.stream()` uses a per-call options-based context: `enhancedOptions._streamDedupContext` (type `StreamGenerationEndContext = { providerEmitted: boolean }`) is created fresh per `sdk.stream()` invocation. Native providers (googleVertex.ts, googleAiStudio.ts) call `markStreamProviderEmittedGenerationEnd(options)` before their emit; the `runStandardStreamRequest` finally block reads `dedupContext.providerEmitted` and skips the orchestration emit when already true. amazonBedrock.ts and ollama.ts have independent stream paths that never enter `runStandardStreamRequest` so they do not need instrumentation. The orchestration emit carries `pipelineAHandled: true` while native provider emits do not.

Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.


🧠 Learnings used
Learnt from: murdore
Repo: juspay/neurolink PR: 0
File: :0-0
Timestamp: 2026-04-17T04:30:25.521Z
Learning: In neurolink (src/lib/neurolink.ts), the `pipelineAHandled: true` flag on `generation:end` events is set ONLY by `finalizeGenerateRequestResult` (the SDK-level wrapper). Native providers (amazonBedrock.ts, ollama.ts, googleAiStudio.ts, googleVertex.ts) emit their own `generation:end` events WITHOUT this flag. The `initializeMetricsListeners()` early-return guard `if (data.pipelineAHandled) return` correctly prevents duplicates: it skips the SDK-level event while allowing native provider events through to Pipeline B.

Learnt from: Yaswanth-2874
Repo: juspay/neurolink PR: 880
File: src/lib/neurolink.ts:6191-6195
Timestamp: 2026-03-17T11:20:55.110Z
Learning: In neurolink stream fallback logging (src/lib/neurolink.ts handleStreamFallback), typical usage is to override both provider and model together, but logs should still accurately reflect one-sided overrides when only model or only provider is set.

Curator P2-4: cost listeners that subscribe to `generation:end` previously
received zero events from `sdk.stream()` calls. The doc described "fires
twice"; the bug on shipped 9.56.x is the opposite direction — `stream()`
emitted `stream:complete` but never `generation:end`, leaving any
listener (cost-tracker, audit log, alerting) with nothing.

Reproduction (all real providers):

  generate / vertex            count=1   PASS
  stream  / vertex             count=0   FAIL  ← bug
  generate / google-ai-studio  count=1   PASS
  stream  / google-ai-studio   count=0   FAIL  ← bug
  generate / litellm           count=1   PASS
  stream  / litellm            count=0   FAIL  ← bug

Fix: in `runStandardStreamRequest`'s `processedStream` generator, emit
`generation:end` exactly once in the `finally` block with the final
stream state (provider, model, content, usage, finishReason, toolsUsed,
prompt, temperature, maxTokens, success, error, pipelineAHandled).
Hoist `resolvedUsage` to the generator scope so it's available in
`finally`. The event payload mirrors `generate()`'s shape so listeners
receive a consistent contract across both APIs.

Backward compatibility: additive — listeners that previously received
zero events on streams now receive one event with the same shape
`generate()` produces.
@murdore
murdore force-pushed the fix/curator-issue-04-stream-generation-end branch from 9c67e68 to 9011edd Compare April 26, 2026 07:18
@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@murdore
murdore merged commit 9bd2cd0 into release Apr 26, 2026
15 checks passed
@murdore
murdore deleted the fix/curator-issue-04-stream-generation-end branch April 26, 2026 07:24
@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 9.59.1 🎉

The release is available on:

Your semantic-release bot 📦🚀

This branch was successfully deployed

1 active deployment
Preview — 9011eddd Deployed Apr 26, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants