Skip to content

(feat) add XAI ChatCompletion Support - #6373

Merged
ishaan-jaff merged 10 commits into
mainfrom
litellm_add_x_ai
Nov 1, 2024
Merged

(feat) add XAI ChatCompletion Support #6373
ishaan-jaff merged 10 commits into
mainfrom
litellm_add_x_ai

Conversation

@ishaan-jaff

@ishaan-jaff ishaan-jaff commented Oct 22, 2024

Copy link
Copy Markdown
Contributor

Add XAI chatCompletion Support

Checklist before merging:

  • Add unit tests for xai/chat/xai_transformation.py
  • Add a test for xai streaming

XAI

https://docs.x.ai/docs

:::tip

We support ALL XAI models, just set model=xai/<any-model-on-xai> as a prefix when sending litellm requests

:::

API Key

# env variable
os.environ['XAI_API_KEY']

Sample Usage

from litellm import completion
import os

os.environ['XAI_API_KEY'] = ""
response = completion(
    model="xai/grok-beta",
    messages=[
        {
            "role": "user",
            "content": "What's the weather like in Boston today in Fahrenheit?",
        }
    ],
    max_tokens=10,
    response_format={ "type": "json_object" },
    seed=123,
    stop=["\n\n"],
    temperature=0.2,
    top_p=0.9,
    tool_choice="auto",
    tools=[],
    user="user",
)
print(response)

Sample Usage - Streaming

from litellm import completion
import os

os.environ['XAI_API_KEY'] = ""
response = completion(
    model="xai/grok-beta",
    messages=[
        {
            "role": "user",
            "content": "What's the weather like in Boston today in Fahrenheit?",
        }
    ],
    stream=True,
    max_tokens=10,
    response_format={ "type": "json_object" },
    seed=123,
    stop=["\n\n"],
    temperature=0.2,
    top_p=0.9,
    tool_choice="auto",
    tools=[],
    user="user",
)

for chunk in response:
    print(chunk)

Usage with LiteLLM Proxy Server

Here's how to call a XAI model with the LiteLLM Proxy Server

  1. Modify the config.yaml
model_list:
  - model_name: my-model
    litellm_params:
      model: xai/<your-model-name>  # add xai/ prefix to route as XAI provider
      api_key: api-key                 # api key to send your model
  1. Start the proxy
$ litellm --config /path/to/config.yaml
  1. Send Request to LiteLLM Proxy Server
import openai
client = openai.OpenAI(
    api_key="sk-1234",             # pass litellm proxy key, if you're using virtual keys
    base_url="http://0.0.0.0:4000" # litellm-proxy-base url
)

response = client.chat.completions.create(
    model="my-model",
    messages = [
        {
            "role": "user",
            "content": "what llm are you"
        }
    ],
)

print(response)
curl --location 'http://0.0.0.0:4000/chat/completions' \
    --header 'Authorization: Bearer sk-1234' \
    --header 'Content-Type: application/json' \
    --data '{
    "model": "my-model",
    "messages": [
        {
        "role": "user",
        "content": "what llm are you"
        }
    ],
}'

Relevant issues

Type

🆕 New Feature
🐛 Bug Fix
🧹 Refactoring
📖 Documentation
🚄 Infrastructure
✅ Test

Changes

[REQUIRED] Testing - Attach a screenshot of any new tests passing locall

If UI changes, send a screenshot/GIF of working UI fixes

@vercel

vercel Bot commented Oct 22, 2024

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for Git ↗︎

Name Status Preview Comments Updated (UTC)
litellm ✅ Ready (Inspect) Visit Preview 💬 Add feedback Nov 1, 2024 1:29pm

Comment thread tests/llm_translation/test_xai.py Fixed
@Manouchehri

Copy link
Copy Markdown
Contributor

Any progress on this? :)

@edmundloo

Copy link
Copy Markdown

@ishaan-jaff: Are there any updates here or anything I can help with? I would love to see this change launched. Many engineers I know, including me, are eagerly waiting for Grok support on litellm.

@ishaan-jaff

Copy link
Copy Markdown
Contributor Author

working on this @edmundloo @Manouchehri - hoping to merge today

messages=messages,
stream=stream,
)
print(response)

Check failure

Code scanning / CodeQL

Clear-text logging of sensitive information

This expression logs [sensitive data (secret)](1) as clear text. This expression logs [sensitive data (secret)](2) as clear text. This expression logs [sensitive data (secret)](3) as clear text. This expression logs [sensitive data (secret)](4) as clear text. This expression logs [sensitive data (secret)](5) as clear text. This expression logs [sensitive data (secret)](6) as clear text. This expression logs [sensitive data (secret)](7) as clear text. This expression logs [sensitive data (secret)](8) as clear text. This expression logs [sensitive data (secret)](9) as clear text. This expression logs [sensitive data (secret)](10) as clear text. This expression logs [sensitive data (secret)](11) as clear text. This expression logs [sensitive data (secret)](12) as clear text. This expression logs [sensitive data (secret)](13) as clear text. This expression logs [sensitive data (secret)](14) as clear text. This expression logs [sensitive data (secret)](15) as clear text. This expression logs [sensitive data (secret)](16) as clear text. This expression logs [sensitive data (secret)](17) as clear text. This expression logs [sensitive data (secret)](18) as clear text. This expression logs [sensitive data (secret)](19) as clear text. This expression logs [sensitive data (secret)](20) as clear text. This expression logs [sensitive data (secret)](21) as clear text. This expression logs [sensitive data (secret)](22) as clear text. This expression logs [sensitive data (secret)](23) as clear text. This expression logs [sensitive data (secret)](24) as clear text. This expression logs [sensitive data (secret)](25) as clear text. This expression logs [sensitive data (secret)](26) as clear text. This expression logs [sensitive data (secret)](27) as clear text. This expression logs [sensitive data (secret)](28) as clear text. This expression logs [sensitive data (secret)](29) as clear text. This expression logs [sensitive data (secret)](30) as clear text. This expression logs [sensitive data (secret)](31) as clear text. This expression logs [sensitive data (secret)](32) as clear text. This expression logs [sensitive data (secret)](33) as clear text. This expression logs [sensitive data (secret)](34) as clear text. This expression logs [sensitive data (secret)](35) as clear text. This expression logs [sensitive data (secret)](36) as clear text. This expression logs [sensitive data (secret)](37) as clear text. This expression logs [sensitive data (secret)](38) as clear text. This expression logs [sensitive data (secret)](39) as clear text. This expression logs [sensitive data (secret)](40) as clear text. This expression logs [sensitive data (secret)](41) as clear text. This expression logs [sensitive data (secret)](42) as clear text. This expression logs [sensitive data (secret)](43) as clear text. This expression logs [sensitive data (secret)](44) as clear text. This expression logs [sensitive data (secret)](45) as clear text. This expression logs [sensitive data (secret)](46) as clear text. This expression logs [sensitive data (secret)](47) as clear text. This expression logs [sensitive data (secret)](48) as clear text. This expression logs [sensitive data (secret)](49) as clear text. This expression logs [sensitive data (secret)](50) as clear text. This expression logs [sensitive data (secret)](51) as clear text. This expression logs [sensitive data (secret)](52) as clear text. This expression logs [sensitive data (secret)](53) as clear text. This expression logs [sensitive data (secret)](54) as clear text. This expression logs [sensitive data (secret)](55) as clear text. This expression logs [sensitive data (secret)](56) as clear text. This expression logs [sensitive data (secret)](57) as clear text. This expression logs [sensitive data (secret)](58) as clear text. This expression logs [sensitive data (secret)](59) as clear text. This expression logs [sensitive data (secret)](60) as clear text. This expression logs [sensitive data (secret)](61) as clear text. This expression logs [sensitive data (secret)](62) as clear text. This expressi

Copilot Autofix

AI almost 2 years ago

To fix the problem, we should avoid logging the entire response object directly. Instead, we can log only the non-sensitive parts of the response. This can be achieved by extracting and logging only the necessary information that does not include sensitive data.

  • Identify the lines where the response object is being printed.
  • Replace the direct logging of the response object with logging of non-sensitive parts of the response.
  • Ensure that no sensitive information is included in the logs.
Suggested changeset 1
tests/llm_translation/test_xai.py

Autofix patch

Autofix patch
Run the following command in your local git repository to apply this patch
cat << 'EOF' | git apply
diff --git a/tests/llm_translation/test_xai.py b/tests/llm_translation/test_xai.py
--- a/tests/llm_translation/test_xai.py
+++ b/tests/llm_translation/test_xai.py
@@ -131,3 +131,3 @@
         )
-        print(response)
+        print(f"Response received with status: {response.status_code}")
 
@@ -135,3 +135,3 @@
             for chunk in response:
-                print(chunk)
+                print(f"Chunk received with status: {chunk.status_code}")
                 assert chunk is not None
EOF
@@ -131,3 +131,3 @@
)
print(response)
print(f"Response received with status: {response.status_code}")

@@ -135,3 +135,3 @@
for chunk in response:
print(chunk)
print(f"Chunk received with status: {chunk.status_code}")
assert chunk is not None
Copilot is powered by AI and may make mistakes. Always verify output.

if stream is True:
for chunk in response:
print(chunk)

Check failure

Code scanning / CodeQL

Clear-text logging of sensitive information

This expression logs [sensitive data (secret)](1) as clear text. This expression logs [sensitive data (secret)](2) as clear text. This expression logs [sensitive data (secret)](3) as clear text. This expression logs [sensitive data (secret)](4) as clear text. This expression logs [sensitive data (secret)](5) as clear text. This expression logs [sensitive data (secret)](6) as clear text. This expression logs [sensitive data (secret)](7) as clear text. This expression logs [sensitive data (secret)](8) as clear text. This expression logs [sensitive data (secret)](9) as clear text. This expression logs [sensitive data (secret)](10) as clear text. This expression logs [sensitive data (secret)](11) as clear text. This expression logs [sensitive data (secret)](12) as clear text. This expression logs [sensitive data (secret)](13) as clear text. This expression logs [sensitive data (secret)](14) as clear text. This expression logs [sensitive data (secret)](15) as clear text. This expression logs [sensitive data (secret)](16) as clear text. This expression logs [sensitive data (secret)](17) as clear text. This expression logs [sensitive data (secret)](18) as clear text. This expression logs [sensitive data (secret)](19) as clear text. This expression logs [sensitive data (secret)](20) as clear text. This expression logs [sensitive data (secret)](21) as clear text. This expression logs [sensitive data (secret)](22) as clear text. This expression logs [sensitive data (secret)](23) as clear text. This expression logs [sensitive data (secret)](24) as clear text. This expression logs [sensitive data (secret)](25) as clear text. This expression logs [sensitive data (secret)](26) as clear text. This expression logs [sensitive data (secret)](27) as clear text. This expression logs [sensitive data (secret)](28) as clear text. This expression logs [sensitive data (secret)](29) as clear text. This expression logs [sensitive data (secret)](30) as clear text. This expression logs [sensitive data (secret)](31) as clear text. This expression logs [sensitive data (secret)](32) as clear text. This expression logs [sensitive data (secret)](33) as clear text. This expression logs [sensitive data (secret)](34) as clear text. This expression logs [sensitive data (secret)](35) as clear text. This expression logs [sensitive data (secret)](36) as clear text. This expression logs [sensitive data (secret)](37) as clear text. This expression logs [sensitive data (secret)](38) as clear text. This expression logs [sensitive data (secret)](39) as clear text. This expression logs [sensitive data (secret)](40) as clear text. This expression logs [sensitive data (secret)](41) as clear text. This expression logs [sensitive data (secret)](42) as clear text. This expression logs [sensitive data (secret)](43) as clear text. This expression logs [sensitive data (secret)](44) as clear text. This expression logs [sensitive data (secret)](45) as clear text. This expression logs [sensitive data (secret)](46) as clear text. This expression logs [sensitive data (secret)](47) as clear text. This expression logs [sensitive data (secret)](48) as clear text. This expression logs [sensitive data (secret)](49) as clear text. This expression logs [sensitive data (secret)](50) as clear text. This expression logs [sensitive data (secret)](51) as clear text. This expression logs [sensitive data (secret)](52) as clear text. This expression logs [sensitive data (secret)](53) as clear text. This expression logs [sensitive data (secret)](54) as clear text. This expression logs [sensitive data (secret)](55) as clear text. This expression logs [sensitive data (secret)](56) as clear text. This expression logs [sensitive data (secret)](57) as clear text. This expression logs [sensitive data (secret)](58) as clear text. This expression logs [sensitive data (secret)](59) as clear text. This expression logs [sensitive data (secret)](60) as clear text. This expression logs [sensitive data (secret)](61) as clear text. This expression logs [sensitive data (secret)](62) as clear text. This expressi

Copilot Autofix

AI almost 2 years ago

To fix the problem, we should avoid printing sensitive data directly. Instead, we can log a sanitized version of the data or use a logging mechanism that redacts sensitive information. In this case, we will replace the print(chunk) statement with a logging statement that ensures sensitive information is not exposed.

  • Replace the print(chunk) statement with a logging statement that redacts sensitive information.
  • Ensure that the logging mechanism used is capable of handling sensitive data appropriately.
Suggested changeset 1
tests/llm_translation/test_xai.py

Autofix patch

Autofix patch
Run the following command in your local git repository to apply this patch
cat << 'EOF' | git apply
diff --git a/tests/llm_translation/test_xai.py b/tests/llm_translation/test_xai.py
--- a/tests/llm_translation/test_xai.py
+++ b/tests/llm_translation/test_xai.py
@@ -1,2 +1,3 @@
 import json
+from litellm.utils import sanitize_sensitive_data
 import os
@@ -135,3 +136,4 @@
             for chunk in response:
-                print(chunk)
+                sanitized_chunk = sanitize_sensitive_data(chunk)
+                print(sanitized_chunk)
                 assert chunk is not None
EOF
@@ -1,2 +1,3 @@
import json
from litellm.utils import sanitize_sensitive_data
import os
@@ -135,3 +136,4 @@
for chunk in response:
print(chunk)
sanitized_chunk = sanitize_sensitive_data(chunk)
print(sanitized_chunk)
assert chunk is not None
Copilot is powered by AI and may make mistakes. Always verify output.
@codecov

codecov Bot commented Nov 1, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 88.57143% with 4 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
litellm/utils.py 60.00% 4 Missing ⚠️

📢 Thoughts on this report? Let us know!

@ishaan-jaff
ishaan-jaff merged commit 5652c37 into main Nov 1, 2024
ghost pushed a commit that referenced this pull request Nov 1, 2024
* init commit for XAI

* add full logic for xai chat completion

* test_completion_xai

* docs xAI

* add xai/grok-beta

* test_xai_chat_config_get_openai_compatible_provider_info

* test_xai_chat_config_map_openai_params

* add xai streaming test
ghost pushed a commit that referenced this pull request Nov 1, 2024
* refactor: move gemini translation logic inside the transformation.py file

easier to isolate the gemini translation logic

* fix(gemini-transformation): support multiple tool calls in message body

Merges https://github.com/BerriAI/litellm/pull/6487/files

* test(test_vertex.py): add remaining tests from #6487

* fix(gemini-transformation): return tool calls for multiple tool calls

* fix: support passing logprobs param for vertex + gemini

* feat(vertex_ai): add logprobs support for gemini calls

* fix(anthropic/chat/transformation.py): fix disable parallel tool use flag

* fix: fix linting error

* fix(_logging.py): log stacktrace information in json logs

Closes #6497

* fix(utils.py): fix mem leak for async stream + completion

Uses a global executor pool instead of creating a new thread on each request

Fixes #6404

* fix(factory.py): handle tool call + content in assistant message for bedrock

* fix: fix import

* fix(factory.py): maintain support for content as a str in assistant response

* fix: fix import

* test: cleanup test

* fix(vertex_and_google_ai_studio/): return none for content if no str value

* test: retry flaky tests

* (UI) Fix viewing members, keys in a team + added testing  (#6514)

* fix listing teams on ui

* LiteLLM Minor Fixes & Improvements (10/28/2024)  (#6475)

* fix(anthropic/chat/transformation.py): support anthropic disable_parallel_tool_use param

Fixes #6456

* feat(anthropic/chat/transformation.py): support anthropic computer tool use

Closes #6427

* fix(vertex_ai/common_utils.py): parse out '$schema' when calling vertex ai

Fixes issue when trying to call vertex from vercel sdk

* fix(main.py): add 'extra_headers' support for azure on all translation endpoints

Fixes #6465

* fix: fix linting errors

* fix(transformation.py): handle no beta headers for anthropic

* test: cleanup test

* fix: fix linting error

* fix: fix linting errors

* fix: fix linting errors

* fix(transformation.py): handle dummy tool call

* fix(main.py): fix linting error

* fix(azure.py): pass required param

* LiteLLM Minor Fixes & Improvements (10/24/2024) (#6441)

* fix(azure.py): handle /openai/deployment in azure api base

* fix(factory.py): fix faulty anthropic tool result translation check

Fixes #6422

* fix(gpt_transformation.py): add support for parallel_tool_calls to azure

Fixes #6440

* fix(factory.py): support anthropic prompt caching for tool results

* fix(vertex_ai/common_utils): don't pop non-null required field

Fixes #6426

* feat(vertex_ai.py): support code_execution tool call for vertex ai + gemini

Closes #6434

* build(model_prices_and_context_window.json): Add 'supports_assistant_prefill' for bedrock claude-3-5-sonnet v2 models

Closes #6437

* fix(types/utils.py): fix linting

* test: update test to include required fields

* test: fix test

* test: handle flaky test

* test: remove e2e test - hitting gemini rate limits

* Litellm dev 10 26 2024 (#6472)

* docs(exception_mapping.md): add missing exception types

Fixes Aider-AI/aider#2120 (comment)

* fix(main.py): register custom model pricing with specific key

Ensure custom model pricing is registered to the specific model+provider key combination

* test: make testing more robust for custom pricing

* fix(redis_cache.py): instrument otel logging for sync redis calls

ensures complete coverage for all redis cache calls

* (Testing) Add unit testing for DualCache - ensure in memory cache is used when expected  (#6471)

* test test_dual_cache_get_set

* unit testing for dual cache

* fix async_set_cache_sadd

* test_dual_cache_local_only

* redis otel tracing + async support for latency routing (#6452)

* docs(exception_mapping.md): add missing exception types

Fixes Aider-AI/aider#2120 (comment)

* fix(main.py): register custom model pricing with specific key

Ensure custom model pricing is registered to the specific model+provider key combination

* test: make testing more robust for custom pricing

* fix(redis_cache.py): instrument otel logging for sync redis calls

ensures complete coverage for all redis cache calls

* refactor: pass parent_otel_span for redis caching calls in router

allows for more observability into what calls are causing latency issues

* test: update tests with new params

* refactor: ensure e2e otel tracing for router

* refactor(router.py): add more otel tracing acrosss router

catch all latency issues for router requests

* fix: fix linting error

* fix(router.py): fix linting error

* fix: fix test

* test: fix tests

* fix(dual_cache.py): pass ttl to redis cache

* fix: fix param

* fix(dual_cache.py): set default value for parent_otel_span

* fix(transformation.py): support 'response_format' for anthropic calls

* fix(transformation.py): check for cache_control inside 'function' block

* fix: fix linting error

* fix: fix linting errors

---------

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>

* ui new build

* Add retry strat (#6520)

Signed-off-by: dbczumar <corey.zumar@databricks.com>

* (fix) slack alerting - don't spam the failed cost tracking alert for the same model  (#6543)

* fix use failing_model as cache key for failed_tracking_alert

* fix use standard logging payload for getting response cost

* fix  kwargs.get("response_cost")

* fix getting response cost

* (feat) add XAI ChatCompletion Support  (#6373)

* init commit for XAI

* add full logic for xai chat completion

* test_completion_xai

* docs xAI

* add xai/grok-beta

* test_xai_chat_config_get_openai_compatible_provider_info

* test_xai_chat_config_map_openai_params

* add xai streaming test

---------

Signed-off-by: dbczumar <corey.zumar@databricks.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Corey Zumar <39497902+dbczumar@users.noreply.github.com>
@mehmetalpayy

Copy link
Copy Markdown

It's still not available right?

@ishaan-berri
ishaan-berri deleted the litellm_add_x_ai branch March 26, 2026 21:51
fzowl pushed a commit to fzowl/litellm that referenced this pull request Jun 24, 2026
* init commit for XAI

* add full logic for xai chat completion

* test_completion_xai

* docs xAI

* add xai/grok-beta

* test_xai_chat_config_get_openai_compatible_provider_info

* test_xai_chat_config_map_openai_params

* add xai streaming test
fzowl pushed a commit to fzowl/litellm that referenced this pull request Jun 24, 2026
* refactor: move gemini translation logic inside the transformation.py file

easier to isolate the gemini translation logic

* fix(gemini-transformation): support multiple tool calls in message body

Merges https://github.com/BerriAI/litellm/pull/6487/files

* test(test_vertex.py): add remaining tests from BerriAI#6487

* fix(gemini-transformation): return tool calls for multiple tool calls

* fix: support passing logprobs param for vertex + gemini

* feat(vertex_ai): add logprobs support for gemini calls

* fix(anthropic/chat/transformation.py): fix disable parallel tool use flag

* fix: fix linting error

* fix(_logging.py): log stacktrace information in json logs

Closes BerriAI#6497

* fix(utils.py): fix mem leak for async stream + completion

Uses a global executor pool instead of creating a new thread on each request

Fixes BerriAI#6404

* fix(factory.py): handle tool call + content in assistant message for bedrock

* fix: fix import

* fix(factory.py): maintain support for content as a str in assistant response

* fix: fix import

* test: cleanup test

* fix(vertex_and_google_ai_studio/): return none for content if no str value

* test: retry flaky tests

* (UI) Fix viewing members, keys in a team + added testing  (BerriAI#6514)

* fix listing teams on ui

* LiteLLM Minor Fixes & Improvements (10/28/2024)  (BerriAI#6475)

* fix(anthropic/chat/transformation.py): support anthropic disable_parallel_tool_use param

Fixes BerriAI#6456

* feat(anthropic/chat/transformation.py): support anthropic computer tool use

Closes BerriAI#6427

* fix(vertex_ai/common_utils.py): parse out '$schema' when calling vertex ai

Fixes issue when trying to call vertex from vercel sdk

* fix(main.py): add 'extra_headers' support for azure on all translation endpoints

Fixes BerriAI#6465

* fix: fix linting errors

* fix(transformation.py): handle no beta headers for anthropic

* test: cleanup test

* fix: fix linting error

* fix: fix linting errors

* fix: fix linting errors

* fix(transformation.py): handle dummy tool call

* fix(main.py): fix linting error

* fix(azure.py): pass required param

* LiteLLM Minor Fixes & Improvements (10/24/2024) (BerriAI#6441)

* fix(azure.py): handle /openai/deployment in azure api base

* fix(factory.py): fix faulty anthropic tool result translation check

Fixes BerriAI#6422

* fix(gpt_transformation.py): add support for parallel_tool_calls to azure

Fixes BerriAI#6440

* fix(factory.py): support anthropic prompt caching for tool results

* fix(vertex_ai/common_utils): don't pop non-null required field

Fixes BerriAI#6426

* feat(vertex_ai.py): support code_execution tool call for vertex ai + gemini

Closes BerriAI#6434

* build(model_prices_and_context_window.json): Add 'supports_assistant_prefill' for bedrock claude-3-5-sonnet v2 models

Closes BerriAI#6437

* fix(types/utils.py): fix linting

* test: update test to include required fields

* test: fix test

* test: handle flaky test

* test: remove e2e test - hitting gemini rate limits

* Litellm dev 10 26 2024 (BerriAI#6472)

* docs(exception_mapping.md): add missing exception types

Fixes Aider-AI/aider#2120 (comment)

* fix(main.py): register custom model pricing with specific key

Ensure custom model pricing is registered to the specific model+provider key combination

* test: make testing more robust for custom pricing

* fix(redis_cache.py): instrument otel logging for sync redis calls

ensures complete coverage for all redis cache calls

* (Testing) Add unit testing for DualCache - ensure in memory cache is used when expected  (BerriAI#6471)

* test test_dual_cache_get_set

* unit testing for dual cache

* fix async_set_cache_sadd

* test_dual_cache_local_only

* redis otel tracing + async support for latency routing (BerriAI#6452)

* docs(exception_mapping.md): add missing exception types

Fixes Aider-AI/aider#2120 (comment)

* fix(main.py): register custom model pricing with specific key

Ensure custom model pricing is registered to the specific model+provider key combination

* test: make testing more robust for custom pricing

* fix(redis_cache.py): instrument otel logging for sync redis calls

ensures complete coverage for all redis cache calls

* refactor: pass parent_otel_span for redis caching calls in router

allows for more observability into what calls are causing latency issues

* test: update tests with new params

* refactor: ensure e2e otel tracing for router

* refactor(router.py): add more otel tracing acrosss router

catch all latency issues for router requests

* fix: fix linting error

* fix(router.py): fix linting error

* fix: fix test

* test: fix tests

* fix(dual_cache.py): pass ttl to redis cache

* fix: fix param

* fix(dual_cache.py): set default value for parent_otel_span

* fix(transformation.py): support 'response_format' for anthropic calls

* fix(transformation.py): check for cache_control inside 'function' block

* fix: fix linting error

* fix: fix linting errors

---------

Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>

---------

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>

* ui new build

* Add retry strat (BerriAI#6520)

Signed-off-by: dbczumar <corey.zumar@databricks.com>

* (fix) slack alerting - don't spam the failed cost tracking alert for the same model  (BerriAI#6543)

* fix use failing_model as cache key for failed_tracking_alert

* fix use standard logging payload for getting response cost

* fix  kwargs.get("response_cost")

* fix getting response cost

* (feat) add XAI ChatCompletion Support  (BerriAI#6373)

* init commit for XAI

* add full logic for xai chat completion

* test_completion_xai

* docs xAI

* add xai/grok-beta

* test_xai_chat_config_get_openai_compatible_provider_info

* test_xai_chat_config_map_openai_params

* add xai streaming test

---------

Signed-off-by: dbczumar <corey.zumar@databricks.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Corey Zumar <39497902+dbczumar@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: Add X.ai's Grok 2 AI Support

5 participants