Repository navigation
feat(proxy): add TypeSafe AI Jev evaluate passthrough with registry-priced spend tracking #41607
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We鈥檒l occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
2dc9697
78eb92c
9470aa4
7fca7fa
4f3b90b
b2ef8da
dcbb773
1feaa48
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -96,6 +96,7 @@ | |
| "/langfuse/", | ||
| "/vllm/", | ||
| "/mistral/", | ||
| "/typesafe/", | ||
| "/nvidia_nim/", | ||
| "/groq/", | ||
| "/voyage/", | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -525,6 +525,42 @@ async def mistral_proxy_route( | |
| return received_value | ||
|
|
||
|
|
||
| @router.api_route( | ||
| "/typesafe/{endpoint:path}", | ||
| methods=["GET", "POST"], # mutable-ok: FastAPI route metadata requires a list | ||
| tags=["TypeSafe AI Pass-through", "pass-through"], # mutable-ok: FastAPI route metadata requires a list | ||
| ) | ||
| async def typesafe_proxy_route( | ||
| endpoint: str, | ||
| request: Request, | ||
| fastapi_response: Response, | ||
| user_api_key_dict: Annotated[UserAPIKeyAuth, Depends(user_api_key_auth)], | ||
| ): | ||
| """[Docs](https://docs.litellm.ai/docs/pass_through/typesafe)""" | ||
| base_target_url: Final = get_secret_str("TYPESAFE_API_BASE") or "https://api.typesafe.ai" | ||
| encoded_endpoint: Final = httpx.URL(endpoint).path | ||
| normalized_endpoint: Final = encoded_endpoint if encoded_endpoint.startswith("/") else f"/{encoded_endpoint}" | ||
| base_url: Final = httpx.URL(base_target_url) | ||
| updated_url: Final = base_url.copy_with( | ||
| path=HttpPassThroughEndpointHelpers.join_base_and_endpoint_path(base_url, normalized_endpoint), | ||
| ) | ||
| typesafe_api_key: Final = passthrough_endpoint_router.get_credentials( | ||
| custom_llm_provider="typesafe", | ||
| region_name=None, | ||
| ) | ||
| endpoint_func: Final = create_pass_through_route( | ||
| endpoint=endpoint, | ||
| target=str(updated_url), | ||
| custom_headers={ # mutable-ok: pass-through request headers require a mutable mapping | ||
| "Authorization": f"Bearer {typesafe_api_key}", | ||
| "Content-Type": "application/json", | ||
| }, | ||
| custom_llm_provider="typesafe", | ||
| is_streaming_request=False, | ||
| ) | ||
| return await endpoint_func(request, fastapi_response, user_api_key_dict) | ||
|
Comment on lines
+528
to
+561
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This route and its logging dispatch add TypeSafe-specific behavior outside Rule Used: What: Avoid writing provider-specific code outside of the llms/ directory. Why: This practice ensures better maintainability and reduces complexity over time. Good: ```python # Handle provider-specific logic within llms/vertex_ai/transformation.py ... (source) Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. TypeSafe pass-through routes intentionally use the shared proxy passthrough layer; moving dispatch into llms would break established provider passthrough architecture.
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. That鈥檚 fair. The route delegates to the shared passthrough machinery, and the TypeSafe-specific logging handler is placed alongside the existing provider passthrough handlers under
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Shared passthrough routing and success handlers intentionally live under proxy; provider-specific completion adapters cannot replace this endpoint forwarding and spend-log integration. |
||
|
|
||
|
|
||
| @router.api_route( | ||
| "/milvus/{endpoint:path}", | ||
| methods=["GET", "POST", "PUT", "DELETE", "PATCH"], | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,117 @@ | ||
| from collections.abc import Mapping | ||
| from datetime import datetime | ||
| from typing import Final | ||
|
|
||
| import httpx | ||
| from pydantic import BaseModel, TypeAdapter, ValidationError | ||
|
|
||
| import litellm | ||
| from litellm.litellm_core_utils.litellm_logging import Logging as LiteLLMLoggingObj | ||
| from litellm.litellm_core_utils.litellm_logging import ( | ||
| get_standard_logging_object_payload, # pyright: ignore[reportUnknownVariableType] # legacy helper has an untyped signature | ||
| ) | ||
| from litellm.proxy._types import PassThroughEndpointLoggingTypedDict | ||
| from litellm.types.utils import ModelResponse, StandardPassThroughResponseObject, Usage | ||
|
|
||
|
|
||
| class _TypeSafeUsage(BaseModel): | ||
| input_tokens: int = 0 | ||
| output_tokens: int = 0 | ||
|
|
||
|
|
||
| class _TypeSafeResponse(BaseModel): | ||
| model: str | None = None | ||
| usage: _TypeSafeUsage | None = None | ||
|
|
||
|
|
||
| class _RegistryPricing(BaseModel): | ||
| input_cost_per_token: float = 0.0 | ||
| output_cost_per_token: float = 0.0 | ||
|
|
||
|
|
||
| _TYPESAFE_RESPONSE_ADAPTER: Final = TypeAdapter(_TypeSafeResponse) | ||
| _REGISTRY_PRICING_ADAPTER: Final = TypeAdapter(_RegistryPricing) | ||
|
|
||
|
|
||
| def _parse_typesafe_response(response_body: Mapping[str, object]) -> _TypeSafeResponse: | ||
| try: | ||
| return _TYPESAFE_RESPONSE_ADAPTER.validate_python(response_body) | ||
| except ValidationError: | ||
| return _TypeSafeResponse() | ||
|
|
||
|
|
||
| def _pricing_for(model_keys: tuple[str, ...]) -> _RegistryPricing: | ||
| for model_key in model_keys: | ||
| if model_key not in litellm.model_cost: # pyright: ignore[reportUnknownMemberType] # registry is dynamically typed | ||
| continue | ||
| try: | ||
| return _REGISTRY_PRICING_ADAPTER.validate_python( | ||
| litellm.model_cost[model_key] # pyright: ignore[reportUnknownMemberType] # registry is dynamically typed | ||
| ) | ||
| except ValidationError: | ||
| continue | ||
| return _RegistryPricing() | ||
|
|
||
|
|
||
| class TypeSafePassthroughLoggingHandler: | ||
| @staticmethod | ||
| def typesafe_passthrough_handler( | ||
| httpx_response: httpx.Response, | ||
| response_body: Mapping[str, object], | ||
| logging_obj: LiteLLMLoggingObj, | ||
| url_route: str, | ||
| result: str, | ||
| start_time: datetime, | ||
| end_time: datetime, | ||
| cache_hit: bool, | ||
| request_body: Mapping[str, object], | ||
| **kwargs: object, | ||
| ) -> PassThroughEndpointLoggingTypedDict: | ||
| response: Final = _parse_typesafe_response(response_body) | ||
| response_model: Final = response.model | ||
| request_model_value: Final = request_body.get("model") | ||
| request_model: Final = request_model_value if isinstance(request_model_value, str) else None | ||
| logged_model: Final = response_model or request_model or "unknown" | ||
| model_name: Final = f"typesafe/{logged_model}" | ||
| usage: Final = response.usage or _TypeSafeUsage() | ||
| input_tokens: Final = usage.input_tokens | ||
| output_tokens: Final = usage.output_tokens | ||
| candidate_model_keys: Final = tuple( | ||
| f"typesafe/{model}" for model in (response_model, request_model) if model is not None | ||
| ) | ||
| pricing: Final = _pricing_for(candidate_model_keys) | ||
|
cursor[bot] marked this conversation as resolved.
|
||
| response_cost: Final = ( | ||
| input_tokens * pricing.input_cost_per_token + output_tokens * pricing.output_cost_per_token | ||
| ) | ||
| usage_object: Final = Usage( | ||
| prompt_tokens=input_tokens, | ||
| completion_tokens=output_tokens, | ||
| total_tokens=input_tokens + output_tokens, | ||
| ) | ||
| updated_kwargs: Final = { # mutable-ok: pass-through logging contract requires mutable kwargs | ||
| **kwargs, | ||
| "model": model_name, | ||
| "custom_llm_provider": "typesafe", | ||
| "response_cost": response_cost, | ||
| "combined_usage_object": usage_object, | ||
| } | ||
| logging_obj.model_call_details.update( | ||
| model=model_name, | ||
| custom_llm_provider="typesafe", | ||
| response_cost=response_cost, | ||
| ) | ||
| standard_logging_object: Final = get_standard_logging_object_payload( | ||
| kwargs=updated_kwargs, | ||
| init_response_obj=ModelResponse(model=model_name, usage=usage_object), | ||
| start_time=start_time, | ||
| end_time=end_time, | ||
| logging_obj=logging_obj, | ||
| status="success", | ||
| ) | ||
| return { # mutable-ok: pass-through logging contract requires mutable result | ||
| "result": StandardPassThroughResponseObject(response=result), | ||
| "kwargs": { # mutable-ok: pass-through logging contract requires mutable kwargs | ||
| **updated_kwargs, | ||
| "standard_logging_object": standard_logging_object, | ||
| }, | ||
| } | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -427,6 +427,7 @@ | |
| "chat", | ||
| "completion", | ||
| "embedding", | ||
| "evaluation", | ||
| "guardrail", | ||
| "image_edit", | ||
| "image_generation", | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
For TypeSafe JSON POST requests such as
/typesafe/v1/systemone?trace=yes, this URL construction no longer includes the incoming query string. The shared JSON passthrough path only reads query parameters from the request body'squery_paramsfield, so normal URL query parameters are omitted from the upstream request. The updated test hides this behavior by mergingrequest.query_paramsinside its fake endpoint rather than exercising the production forwarding path.There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The shared passthrough forwards request.query_params itself (pass_through_endpoints.py line 994, then params= on the JSON POST build_request), so removing them from the target avoids duplication.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
You're right. The shared JSON passthrough adds
request.query_paramsas theparams=argument when building the upstream request, so keeping the query string out oftargetavoids duplication while still forwarding it. The test'scopy_merge_params()models that shared behavior; it is not masking a production omission. This comment is incorrect and can be dismissed.