Skip to content

feat(voyage): add voyage-context-4 + voyage-4 series models - #3

Closed
fzowl wants to merge 1 commit into
mainfrom
feat/voyage-context-4
Closed

feat(voyage): add voyage-context-4 + voyage-4 series models#3
fzowl wants to merge 1 commit into
mainfrom
feat/voyage-context-4

Conversation

@fzowl

@fzowl fzowl commented Jul 26, 2026

Copy link
Copy Markdown
Owner

Summary

Adds VoyageAI's latest models to the cost map and hardens contextual embedding input handling.

Models added (model_prices_and_context_window_backup.json)

  • voyage/voyage-context-4 — contextual, 120K ctx, $0.12/M
  • voyage/voyage-4, voyage/voyage-4-large, voyage/voyage-4-lite — 32K ctx, $0.06 / $0.12 / $0.02 /M
  • voyage/voyage-multimodal-3.5 — 32K ctx, $0.12/M

Contextual embeddings input fix

VoyageContextualEmbeddingConfig now normalizes input to list[list[str]] before calling the contextualized embeddings API. A plain str becomes [[str]] and a flat list[str] becomes a single document [list[str]]; already-nested input is passed through unchanged. Ensures the API is always called with list[str] chunk groups.

Docs

Updated docs/providers/voyage.md — supported models table, contextual section (now voyage-context-4), and model selection guide.

Tests

tests/llm_translation/test_voyage_ai.py — context-4 detection + flat/str input normalization. All 16 pass locally.

Reviewer double-checks

  • Prices/context lengths sourced from VoyageAI pricing + embeddings docs (context-4 total 120K, per-chunk 32K).
  • Flat list[str] normalized as a single document; confirm that matches intended semantics.

- Add voyage-context-4, voyage-4, voyage-4-large, voyage-4-lite, and
  voyage-multimodal-3.5 to the model cost map
- Normalize contextual embedding input to list[list[str]] so the
  contextualized embeddings API is always called with list[str] groups
  (accepts str and flat list[str] inputs)
- Update Voyage provider docs and add tests
@fzowl fzowl closed this Jul 26, 2026
@fzowl
fzowl deleted the feat/voyage-context-4 branch July 26, 2026 21:15
fzowl pushed a commit that referenced this pull request Jul 29, 2026
…api (BerriAI#33110)

* feat(guardrails): support streaming text transformation in generic_guardrail_api

* chore(guardrails): address PR review feedback

* fix(guardrails): fail closed on tool-call and prefix-rewrite leaks in streaming transform

* fix(guardrails): address Bugbot review on streaming transform correctness

* fix(guardrails): coerce holdback in handler for in-process guardrails

* fix(guardrails): harden streaming transform (holdback coercion, tool-call passthrough, n>1 finish_reason)

* test(guardrails): targeted _mode_matches coverage for all guardrail_mode shapes

* fix(guardrails): inspect streamed tool calls and harden incremental_diff edge cases

* test: move ComplianceChecker mode tests to the compliance PR

* fix(guardrails): strip content from tool-call passthrough so streamed text can't bypass the transform

* fix(guardrails): four correctness fixes for incremental_diff streaming path

Four bug fixes on top of the OSS PR's incremental_diff streaming text
transformation, all inside the incremental_diff code paths only. No
existing block_only, non-streaming, or pre_call behavior is touched.

Fix #1 — Mixed content+tool_call finish_reason ordering
  _tool_call_passthrough_chunk now takes an optional finish_reason_per_choice
  map. For a choice carrying both delta.content and delta.tool_calls,
  finish_reason is stripped from the passthrough and recorded on the map so
  the final synthetic text chunk delivers it. Without this, SSE-compliant
  clients stopping at finish_reason drop the guardrailed text — defeating
  the redaction the whole feature exists for. (Greptile P1 twice, Veria.)

Fix #2 — Choice index sort in _process_streaming_transform
  indices/texts_to_check were derived from dict insertion order. For n>1
  streams where choice 1 emits before choice 0, guardrail-returned texts
  aligned to the input order mapped back to the wrong choice indices on
  write-back — wrong text goes to wrong choice. Sort raw_by_index.keys()
  up front so realignment is deterministic. (Bugbot Medium.)

Fix #3 — Cross-chunk pre-tool-call text flush
  With default streaming_sampling_rate=5, text chunks followed by a pure
  tool-call chunk carrying finish_reason='tool_calls' would emit the
  passthrough with finish_reason before any transformed text delta had
  fired. Same failure mode as fix #1 but cross-chunk. Now we flush any
  accumulated text via _round(is_final=False) BEFORE yielding the
  tool-call passthrough. (Greptile P1.)

Fix #4 — Terminator chunk for deferred finish_reason on empty mutated_text
  _build_transform_chunk returned None early when mutated_text_per_choice
  was empty. If a mixed content+tool_call chunk had deferred its
  finish_reason (via fix #1) and the guardrail then suppressed the text
  (empty return), the deferred finish_reason was never delivered. Now on
  is_final=True with empty mutated_text_per_choice, we emit a terminator
  carrying finish_reason per choice from finish_reason_per_choice.
  (Bugbot High.)

Also normalized Optional[X] → X | None across the OSS PR's added surface
via ruff UP045 autofix to keep the strict-rule gate within budget. Pure
mechanical typing style change, no semantic effect.

Regression tests for all four fixes:
- test_mixed_chunk_finish_reason_arrives_after_transformed_text (#1)
- test_text_flush_precedes_tool_call_passthrough (#3)
- test_final_finish_reason_flushed_when_guardrail_suppresses_text (#4)
- test_transform_sends_texts_sorted_by_choice_index (#2)

All fixes reachable only when streaming_transform_mode == 'incremental_diff'
is configured (via _run_incremental_transform_stream) or when a
StreamTransformSink is present (via _process_streaming_transform). Verified
scope-clean: no changes to block_only, non-streaming, pre_call, moderation,
or sibling guardrails.

---------

Co-authored-by: Marton Schneider <marton@schneider.co.nl>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant