Repository navigation
feat(batches): support Mistral files/batches and per-page OCR batch cost tracking - #40484
mubashir1osmani wants to merge 16 commits into
Conversation
…ost tracking Adds MistralFilesConfig and MistralBatchesConfig so Mistral can be used as a Files and Batches provider through the shared BaseLLMHTTPHandler path, the same way Bedrock plugs in. /v1/ocr is now an accepted batch endpoint, and completed OCR batches are billed per page (ocr_cost_per_page_batches, half the synchronous rate) instead of per token. Resolves BerriAI#29914
…ider
GET /v1/files/{id} for an id encoded with a non-OpenAI deployment forwarded
the deployment credentials but let custom_llm_provider default to openai, so
a Mistral file was fetched from api.openai.com with the Mistral key and 401'd.
Delete and content already passed the provider through; retrieve now does too.
… files/batches configs
|
PR overviewAll previously flagged issues have been addressed. No open security concerns remain on this pull request. Security reviewNo open security issues remain on this pull request. Fixed/addressed: 3 · PR risk: 0/10 |
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
mateo-berri
left a comment
There was a problem hiding this comment.
Requesting changes only for the two required reds and the test banners, all this PR's own; diff and live QA (body) look right
…tellm_mistral_ocr_batches
… file and batch credentials Files and batches routes take their model from a header, query param or a model-encoded resource id, which the auth layer never sees, so any key could name any deployment and act on that provider account with its server-side key. Every caller-supplied model now goes through can_key_call_resolved_model before deployment credentials are resolved, covering file create/retrieve/content/ delete/list, batch create/retrieve/list/cancel, and vector store files.
… routes Unified ids carry the deployment model inside the id, so a restricted key could create, retrieve or cancel a batch on a deployment it is not granted. The model parsed from a unified id now goes through the same grant check as header, query and model-encoded id sources before the router is called.
6c750a8 to
bae731d
Compare
…atch retrieves Provider batch configs that build their own request URL (Mistral, Bedrock) hand pre_call api_base=None, and mask_api_base_credentials raised TypeError on it, so every such retrieve logged a non-blocking LoggingError and lost its pre-call logging.
|
bugbot run |
|
@mateo-berri please review |
…hem to batch The proxy runs batch-file validation and guardrails only for purpose=batch, so a purpose such as assistants that was silently rewritten to batch on the way to Mistral let an upload skip both. Only batch, fine-tune and ocr pass through now; anything else is a 400.
…tches # Conflicts: # litellm/batches/batch_utils.py # tests/test_litellm/llms/mistral/ocr/test_mistral_ocr_cost.py
…sage text main (15f2e25) replaced the configurable model-access-denied message with a fixed client message, so match on the stable error type.
get_model_info raises a bare Exception for unmapped models, so BLE001 cannot be narrowed; mark it noqa with the reason to stay within the strict budget.
|
bugbot run |
… logging A pre-signed batch/file request (Mistral, Bedrock) carries its auth header inside the transformed request body, which pre_call logs verbatim into raw_request_typed_dict and raw-request callbacks, leaking the provider key. Mask the nested headers channel before handing the request to pre_call. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…atches # Conflicts: # litellm/batches/batch_utils.py
…ost warnings single-line
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit f3b198c. Configure here.
…cut and skip it for registry-routed vector store models retrieve_batch returned a terminal batch from the DB before checking that the key may use the model encoded in a unified batch id; the grant check now runs right after pre-call processing. The vector store file list helper authorized data["model"] through handle_model_based_routing even when the vector store registry set it server-side and even with no caller, which crashed on a None key; it now authorizes only a caller-supplied hint and resolves credentials directly.
| data=data, | ||
| check_file_id_encoding=False, | ||
| ) | ||
| credentials = get_credentials_for_model(llm_router=llm_router, model_id=model_hint) |
There was a problem hiding this comment.
For a model derived from the vector-store registry, this branch calls get_credentials_for_model without authorizing that model for the caller. Access to a vector store does not guarantee access to its backing deployment: teamless stores are available to every caller, and object permissions can grant store access independently of model grants. A key restricted to other models can therefore use credentials for a deployment it is not allowed to access. The model grant must be checked before resolving its credentials. This violates the repository directive to reject authentication-layer changes that may cause a security incident.
How this was verified: Registry-derived model hints reach the credential resolver without a model-grant check, while vector-store access can be granted independently through teamless stores, object permissions, or team membership.
Rule Used: What: Fail any PR which may contains a security incident on litellm's authentication layer Why: Do not cause security incidents Bad: ```python # Check cache first cache_key = ( f"oidc_userinfo_{token[:20]}" # Use fi... (source)
Knowledge Base Used: Proxy authentication and authorization
There was a problem hiding this comment.
Unchanged from main: #29739 authorizes only caller-supplied hints. Registry-bound stores are admin config gated by vector store permissions. Authorizing them would 403 existing store-permitted teams
There was a problem hiding this comment.
That distinction is valid. should_authorize_model_hint is only true when the model came from ?model/x-litellm-model; a model injected from the admin-controlled vector-store registry is intentionally not treated as a caller-supplied hint. Access is gated earlier by assert_user_can_access_vector_store_id, and adding the model-grant check here would change the established behavior for teams authorized to use that store. Since this matches the existing #29739 boundary and the diff does not introduce a new authorization bypass for registry-bound stores, I’m withdrawing this finding.
1 similar comment
|
Superseded by #41934, an internal-branch copy of this exact tip so CircleCI can run. Your commits and authorship carry over unchanged |
feat(batches): support Mistral files/batches and per-page OCR batch cost tracking (internal copy of #40484)
TLDR
Problem this solves:
/v1/ocris rejected as a batch endpoint, so OCR batches are impossibleHow it solves it:
MistralFilesConfigandMistralBatchesConfigon the shared HTTP handler path/v1/ocras a batchendpointocr_cost_per_page_batchesrateGET /v1/files/{id}sending model-routed non-OpenAI ids to api.openai.comTypeErrorin pre-call logging when a presigned batch retrieve passesapi_base=NoneUser Flow
Before: a developer who wants Mistral's discounted OCR batches cannot get past the upload, so the whole job runs outside LiteLLM's auth and spend tracking
purpose=batch,model=mistral-ocrand a JSONL of/v1/ocrrequests and get 400LiteLLM doesn't support mistral for 'create_file'. Only ['openai', 'azure', 'vertex_ai', 'manus', 'anthropic'] are supported.endpoint: /v1/ocr,model: mistral-ocranyway and get 400LiteLLM doesn't support custom_llm_provider=mistral for 'create_batch'LiteLLM doesn't support mistral for 'file_list'After: the same developer runs the whole OCR batch lifecycle through the gateway and sees per-page spend
purpose=batch,model=mistral-ocrand the same JSONLfile-...idendpoint: /v1/ocr,model: mistral-ocrand get back abatchobject withstatus: validatingstatus: completed, then GET http://localhost:4000/v1/files/{output_file_id}/content to download the OCR resultsclaude-haikusends GET http://localhost:4000/v1/files/{that id} and gets 403The requested model 'mistral-ocr' is not available for this API key, or the model name is invalid, so the Mistral deployment's key is never used on their behalfRelevant issues
Fixes #29914
Affected release
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
/qa verdict: PASS (before = fail, after = succeeds). Two legs, each in its own worktree, venv and Postgres database, booted with
python litellm/proxy/proxy_cli.py --config config.yaml --port <port> --num_workers 2andLITELLM_LOCAL_MODEL_COST_MAP=True, against the real Mistral API with a real key, so the After batch cost real money. Before is the merge base faed57f on :21905, After is 5783a38 on :33468. Both legs use the same config and the same 3-line input, and the flow sends each request exactly onceocr_batch_input.jsonlis 3 lines, each a one-page PDF as a base64 data URI:{"custom_id": "doc-0", "method": "POST", "url": "/v1/ocr", "body": {"document": {"type": "document_url", "document_url": "data:application/pdf;base64,JVBERi0x..."}}}Before (faed57f)
GET /spend/logs/v2over the run's window returned[]at query time and the logs page shows the three calls as $0 Failure rows, nothing billable:After (5783a38)
GET /v1/batches/$BATCH_IDevery 15 s:in_progressat 05:31:31Z,in_progressat 05:31:47Z,completedat 05:32:02Z:GET /spend/logs/v2over the run's window, one line per row, shows the batch billed at $0.006, which is 3 pages at the $0.002 per-page batch rate and half the $0.012 sync price:The logs page shows the same batch row at $0.006000, and its detail view shows 3 successful requests on
mistral-ocrwith cost $0.00600000:Observations from both legs, none of them a defect in what this PR changes:
model,usage,expires_atalways null: PR causes_batch_costrow: pre-existing, leaves alonemistral-ocr, create rowmistral/mistral-ocr-latest: PR causesafile_contentrows, empty model, $0: leaves alonealist_fine_tuning_jobs: pre-existing, leaves aloneGET /v1/files/{id}onlyType
🆕 New Feature
🐛 Bug Fix
Caveats (if any)
Medium
Low
list_batchesandcancel_batchare not wired for Mistral, same as Bedrockmethodandurl, which Mistral ignoresrequest_counts.total: 0until validation finishesmodel,usageandexpires_atnull, and the batch cost row names the model group (mistral-ocr) where the create row names the deployment (mistral/mistral-ocr-latest)purpose: user_data;user_datamaps back ontoocron upload and listassistants,visionandevalsuploads get 400 instead of silently becomingbatchFinal Attestation
Note
Medium Risk
Touches proxy authorization for file/batch credential routing and batch billing logic for OCR; both are security- and spend-sensitive but are covered by new tests.
Overview
Adds Mistral as a first-class Files and Batches provider (create/retrieve jobs, multipart uploads, purpose mapping for
batch/ocr/fine-tune) and extends batch APIs to accept/v1/ocrwithmistralin provider literals.Batch spend accounting now detects OCR output lines via
usage_infoand prices them with newocr_batch_costusingocr_cost_per_page_batches/annotation_cost_per_page_batches(with sync-rate fallback); Mistral OCR models in the cost map gain those fields and/v1/batchinsupported_endpoints.On the proxy, model-encoded file/batch IDs and
x-litellm-modelrouting go throughget_authorized_credentials_for_modelso restricted keys cannot borrow another deployment’s credentials;afile_retrieveforwardscustom_llm_providerfor model-routed retrieves (fixes Mistral ids hitting OpenAI). Logging toleratesapi_base=Noneon presigned batch retrieves and masks auth headers embedded in presigned request bodies before raw-request logs.Reviewed by Cursor Bugbot for commit f3b198c. Bugbot is set up for automated code reviews on this repo. Configure here.