Skip to content

fix(cost_map): mark fireworks_ai/minimax-m3 as vision and pin capabilities verified against live calls - #43390

Open
devin-ai-integration[bot] wants to merge 2 commits into
mainfrom
litellm_fireworks_minimax_m3_vision_pin
Open

devin-ai-integration[bot] wants to merge 2 commits into
mainfrom
litellm_fireworks_minimax_m3_vision_pin

Conversation

@devin-ai-integration

@devin-ai-integration devin-ai-integration Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

How it solves it:

  • Both minimax-m3 cost map rows go back to supports_vision: true
  • New ci_cd/cost_map_pins.json records the live-call evidence for that flag
  • The cost map guard fails any PR that lowers a pinned flag, quoting the evidence
  • The sync bot may not touch the pins file, so only a human with new evidence can

User Flow

Before: a developer whose app sends screenshots to fireworks_ai/minimax-m3 through the proxy gets a 400 from LiteLLM before Fireworks is ever called

  1. The proxy admin adds fireworks_ai/accounts/fireworks/models/minimax-m3 to the config as fireworks-minimax-m3 and starts the proxy
  2. The developer sends POST https://litellm-domain/v1/chat/completions with model fireworks-minimax-m3 and one user message holding a text block plus an image_url block
  3. They get HTTP 400 Fireworks AI model accounts/fireworks/models/minimax-m3 does not support image inputs. Use a Fireworks vision model or remove image_url content blocks. and Fireworks never sees the request
  4. The same body sent to POST https://api.fireworks.ai/inference/v1/chat/completions returns HTTP 200 with the image read correctly, so they know the model can see
  5. GET https://litellm-domain/model/info shows supports_vision: false for the deployment, and a later cost map sync can flip it back to false even after someone fixes it by hand

After: the same request reaches Fireworks and comes back with the image read, and no sync can flip the flag back

  1. The proxy admin adds fireworks_ai/accounts/fireworks/models/minimax-m3 to the config as fireworks-minimax-m3 and starts the proxy
  2. The developer sends POST https://litellm-domain/v1/chat/completions with model fireworks-minimax-m3 and one user message holding a text block plus an image_url block
  3. They get HTTP 200 with the image described in the reply (digit=7, shape=circle, color=red for the probe image), billed at the usual minimax-m3 token prices
  4. GET https://litellm-domain/model/info shows supports_vision: true for the deployment
  5. A PR that lowers that flag again fails the "Cost map guard" check with the live-call evidence and the date it was verified, so the fix stays fixed

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

Shared setup for both legs: a worktree at the named commit with its own venv, proxy_config.yaml as below with the real Fireworks key, booted with LITELLM_LOCAL_MODEL_COST_MAP=True litellm --config proxy_config.yaml --num_workers 2 --port <random free port> so each leg reads its own commit's cost map (the hosted main map still says false until this merges). No database. The probe is a 512x512 PNG with a large 7 and a small red circle, sent the same way through all three endpoints: proxy_request.json carries it as a data:image/png;base64 image_url block, messages_request.json as a base64 image source block, responses_request.json as an input_image data URL. The sync commit for the guard case, fb2bae5a, is the PR tip with supports_vision set back to false on both minimax-m3 rows in both cost map files (a 4-line diff), checked under a bot branch name.

model_list:
  - model_name: fireworks-minimax-m3
    litellm_params:
      model: fireworks_ai/accounts/fireworks/models/minimax-m3
      api_key: os.environ/FIREWORKS_AI_API_KEY
  - model_name: fireworks-kimi-k3
    litellm_params:
      model: fireworks_ai/accounts/fireworks/models/kimi-k3
      api_key: os.environ/FIREWORKS_AI_API_KEY
general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY

Before (b248b1c)

Proxy booted from a worktree at b248b1c7dc on port 35188, 2 uvicorn workers.

Image to fireworks-minimax-m3 through /v1/chat/completions

  1. Run

    $ curl -s -o chat.json -w 'HTTP %{http_code}\n' http://127.0.0.1:35188/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @proxy_request.json && jq -c 'if .choices then {content: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}} else . end' chat.json
    
  2. Observed

    HTTP 400
    {"error":{"message":"litellm.BadRequestError: Fireworks AI model accounts/fireworks/models/minimax-m3 does not support image inputs. Use a Fireworks vision model or remove image_url content blocks.\n\nLiteLLM: model group 'fireworks-minimax-m3' failed with the error above. No fallback was attempted.","type":"invalid_request_error","param":null,"code":"400"}}
    

Image to fireworks-minimax-m3 through /v1/messages

  1. Run

    $ curl -s -o messages.json -w 'HTTP %{http_code}\n' http://127.0.0.1:35188/v1/messages -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @messages_request.json && jq -c 'if .content then {text: [.content[] | select(.type == "text") | .text], stop_reason} else . end' messages.json
    
  2. Observed

    HTTP 400
    {"type":"error","error":{"type":"invalid_request_error","message":"litellm.BadRequestError: Fireworks AI model accounts/fireworks/models/minimax-m3 does not support image inputs. Use a Fireworks vision model or remove image_url content blocks.\n\nLiteLLM: model group 'fireworks-minimax-m3' failed with the error above. No fallback was attempted."}}
    

Image to fireworks-minimax-m3 through /v1/responses

  1. Run

    $ curl -s -o responses.json -w 'HTTP %{http_code}\n' http://127.0.0.1:35188/v1/responses -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @responses_request.json && jq -c 'if .output then {text: [.output[] | select(.type == "message") | .content[0].text], status} else . end' responses.json
    
  2. Observed

    HTTP 200
    {"text":["digit=7, shape=circle, color=red"],"status":"completed"}
    

Control: the same image to fireworks-kimi-k3 through /v1/chat/completions

  1. Run

    $ curl -s -o chat_kimi.json -w 'HTTP %{http_code}\n' http://127.0.0.1:35188/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @proxy_request_kimi.json && jq -c 'if .choices then {content: .choices[0].message.content} else . end' chat_kimi.json
    
  2. Observed

    HTTP 200
    {"content":"digit=7, shape=circle, color=red"}
    

What /model/info reports for both deployments

  1. Run

    $ curl -s http://127.0.0.1:35188/model/info -H "Authorization: Bearer $LITELLM_MASTER_KEY" | jq -c '.data[] | {model_name, supports_vision: .model_info.supports_vision}'
    
  2. Observed

    {"model_name":"fireworks-minimax-m3","supports_vision":false}
    {"model_name":"fireworks-kimi-k3","supports_vision":true}
    

The cost map guard at this commit checking a sync commit that flips minimax-m3 back to no-vision

  1. Run

    $ uv run --no-sync python ci_cd/cost_map_guard.py --base e5cc50e7fc0d30ca7f07c1940488c90de48d5810 --head fb2bae5a77b126da82654aec9c0361929e3bb52b --head-ref litellm_cost_map_sync_x; echo "exit $?"
    
  2. Observed

    cost map guard passed (bot contract enforced)
    exit 0
    

After (e5cc50e)

Proxy booted from a worktree at e5cc50e7fc on port 56311, 2 uvicorn workers.

Image to fireworks-minimax-m3 through /v1/chat/completions

  1. Run

    $ curl -s -o chat.json -w 'HTTP %{http_code}\n' http://127.0.0.1:56311/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @proxy_request.json && jq -c 'if .choices then {content: .choices[0].message.content, usage: {prompt_tokens: .usage.prompt_tokens, completion_tokens: .usage.completion_tokens}} else . end' chat.json
    
  2. Observed

    HTTP 200
    {"content":"digit=7, shape=circle, color=red","usage":{"prompt_tokens":519,"completion_tokens":37}}
    

Image to fireworks-minimax-m3 through /v1/messages

  1. Run

    $ curl -s -o messages.json -w 'HTTP %{http_code}\n' http://127.0.0.1:56311/v1/messages -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @messages_request.json && jq -c 'if .content then {text: [.content[] | select(.type == "text") | .text], stop_reason} else . end' messages.json
    
  2. Observed

    HTTP 200
    {"text":["digit=7, shape=circle, color=red"],"stop_reason":"end_turn"}
    

Image to fireworks-minimax-m3 through /v1/responses

  1. Run

    $ curl -s -o responses.json -w 'HTTP %{http_code}\n' http://127.0.0.1:56311/v1/responses -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @responses_request.json && jq -c 'if .output then {text: [.output[] | select(.type == "message") | .content[0].text], status} else . end' responses.json
    
  2. Observed

    HTTP 200
    {"text":["digit=7, shape=circle, color=red"],"status":"completed"}
    

Control: the same image to fireworks-kimi-k3 through /v1/chat/completions

  1. Run

    $ curl -s -o chat_kimi.json -w 'HTTP %{http_code}\n' http://127.0.0.1:56311/v1/chat/completions -H "Authorization: Bearer $LITELLM_MASTER_KEY" -H 'Content-Type: application/json' -d @proxy_request_kimi.json && jq -c 'if .choices then {content: .choices[0].message.content} else . end' chat_kimi.json
    
  2. Observed

    HTTP 200
    {"content":"digit=7, shape=circle, color=red"}
    

What /model/info reports for both deployments

  1. Run

    $ curl -s http://127.0.0.1:56311/model/info -H "Authorization: Bearer $LITELLM_MASTER_KEY" | jq -c '.data[] | {model_name, supports_vision: .model_info.supports_vision}'
    
  2. Observed

    {"model_name":"fireworks-minimax-m3","supports_vision":true}
    {"model_name":"fireworks-kimi-k3","supports_vision":true}
    

The cost map guard at this commit checking a sync commit that flips minimax-m3 back to no-vision

  1. Run

    $ uv run --no-sync python ci_cd/cost_map_guard.py --base e5cc50e7fc0d30ca7f07c1940488c90de48d5810 --head fb2bae5a77b126da82654aec9c0361929e3bb52b --head-ref litellm_cost_map_sync_x; echo "exit $?"
    
  2. Observed

    cost map guard failed (bot contract enforced):
    - model_prices_and_context_window.json: fireworks_ai/accounts/fireworks/models/minimax-m3.supports_vision must stay true (verified 2026-09-26: POST https://api.fireworks.ai/inference/v1/chat/completions read a 512x512 PNG back as digit=7, shape=circle, color=red while GET https://api.fireworks.ai/inference/v1/models still said supports_image_input false and GET https://api.fireworks.ai/v1/serverless/models?format=nested carried no supportsImageInput on the row); change ci_cd/cost_map_pins.json with new evidence first
    - model_prices_and_context_window.json: fireworks_ai/minimax-m3.supports_vision must stay true (verified 2026-09-26: POST https://api.fireworks.ai/inference/v1/chat/completions read a 512x512 PNG back as digit=7, shape=circle, color=red while GET https://api.fireworks.ai/inference/v1/models still said supports_image_input false and GET https://api.fireworks.ai/v1/serverless/models?format=nested carried no supportsImageInput on the row); change ci_cd/cost_map_pins.json with new evidence first
    exit 1
    

Observations from the run:

  • /v1/responses answered the image at the merge base too; this PR leaves that alone
  • Fireworks' model listing still says supports_image_input: false; this PR leaves that alone

Type

🐛 Bug Fix

Caveats (if any)

Medium

  • The guard workflow runs the base branch's code, so the pins check only bites on PRs opened after this merges
  • Pins are exact-key matches, so a new alias row for the same model is not covered until it is pinned too

Low

  • Fireworks' /inference/v1/models and /v1/serverless/models still say no image input for minimax-m3 (checked 2026-09-26); the pin is what keeps the map from following them
  • The daily model registry audit automation was told about the pins file and not to lower a pinned flag

Live regression check

Verdict: no dependent path regressed. Base (b248b1c7dc) and head were driven through the same proxy config, 2 uvicorn workers, no database, with a forwarding recorder in front of https://api.fireworks.ai on both sides, and the only outbound differences are the three requests the base rejected before calling Fireworks

Breaking

  • None observed. The two requests both sides send (a text-only chat to minimax-m3 and the kimi-k3 image control) are byte-identical on the wire once the timestamp and user-agent are dropped, the header key sets match (accept, authorization, content-type, user-agent), and every request went out exactly once
  • Rig notes: user-agent reads litellm/unknown on base and litellm/1.104.0 on head because the base venv was built with --no-install-project; the first head run had a database attached through a stray .env in the worktree and was redone without it, so both legs are no-database

Backward incompatible

  • supports_vision on both fireworks_ai/minimax-m3 keys goes false -> true and every reader moves with it: litellm.get_model_info() in-process, GET /model/info and GET /model_group/info on the proxy, and the Fireworks pre-flight in litellm/llms/fireworks_ai/chat/transformation.py, which no longer answers 400 for image_url content on /v1/chat/completions (stream and non-stream) and /v1/messages. Those three requests now reach Fireworks carrying exactly one image part each and get billed (519 prompt tokens on the probe); a caller that leaned on the 400 to keep images away from this model now gets a 200
  • Cost per token on both rows is unchanged; the row diff is the one flag, 4 lines across both cost map files
  • The cost map guard's contract for the sync bot: a litellm_cost_map_sync_* PR that lowers a pinned flag fails with the pin's evidence (observed on the minted sync commit, exit 1 with both pin messages), one that edits ci_cd/cost_map_pins.json fails, and a human PR that deletes the pins file fails (Greptile's finding, fixed in e5cc50e7fc, covered by test_main_fails_a_pr_that_deletes_the_pins_file, which runs the script as a subprocess). The workflow runs base-branch code, so all of this bites only on PRs opened after this merges

Regression risk

  • Fireworks could drop image support on minimax-m3 later, and the pin would then keep the map saying true while requests fail with Fireworks' own error instead of LiteLLM's pre-flight 400. The pin carries its verification date so a reader can tell stale from wrong; re-verifying is one live call plus a pins file edit
  • Pins match exact keys, so a new alias row for the same model is not covered until it is pinned too

Dependency graph

Dependent Reached through Status
Fireworks chat pre-flight /v1/chat/completions image, stream and non-stream verified live: base 400, head 200
Fireworks chat pre-flight /v1/messages image verified live: base 400, head 200
Responses bridge /v1/responses image verified live: 200 both sides
Model info readers /model/info, /model_group/info verified live: false -> true
SDK readers litellm.get_model_info in both venvs verified live: false -> true, costs equal
kimi-k3 control /v1/chat/completions image verified live: identical both sides
Text-only chat /v1/chat/completions verified live: identical both sides
Cost map guard the script's CLI as the workflow calls it verified live on the minted sync commit
Cost map guard tests/unit/test_cost_map_guard.py tested: 40 passed
Guard workflow .github/workflows/cost-map-guard.yml unreachable: runs base-branch code until merged
Fireworks transformation tests/unit/llms/fireworks_ai tested: 123 passed on head

Merge-ref: main moved from the merge base b248b1c7dc to e73abe6c72 since the branch point and has not moved since. A merged tree (1216bd66 onto e73abe6c72) served the image chat and messages requests with 200 and reports supports_vision: true, and main's guard passes that merge while failing a follow-up sync that lowers the flag. The one commit since, e5cc50e7fc, touches only the guard script and its tests, which the proxy never imports

Not verified

  • CircleCI: run-ci is not added on weekends, so it runs Monday
  • Streaming variants of /v1/messages and /v1/responses with the image (only chat streaming was run)
  • The Admin UI models page rendering supports_vision for this deployment (the field's type is unchanged, only its value)
  • A proxy reading the hosted cost map after the merge (the hosted map says false until then; both legs used LITELLM_LOCAL_MODEL_COST_MAP=True)
  • An actual GitHub Actions run of the new guard on a bot PR; only the script's CLI was run locally with the workflow's arguments

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

…ities verified against live calls

Fireworks' serverless minimax-m3 answers image inputs, but its model
listings say supports_image_input false, so the last two registry syncs
flipped supports_vision back to false and the Fireworks pre-flight turned
every image request into a 400 without ever calling the provider.

Flip both minimax-m3 rows back to supports_vision true and record the
live-call evidence in ci_cd/cost_map_pins.json. The cost map guard now
fails any PR that lowers a pinned capability, quoting the evidence, and
the sync bot may not touch the pins file.
@devin-ai-integration

devin-ai-integration Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor Author

I'll fix CI failures and address comments from users with write access. I'll skip comments containing "(aside)".

  • Disable automatic comment, CI, and merge conflict monitoring

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Updates model capability metadata and adds validation for pinned overrides.

The PR should not merge until the guard accepts valid older cost-map branches and the new unit-test rule violation is addressed.

Findings

  1. P1 Older branches fail the guard ▶
  2. P2 Unit test runs subprocesses ▶

Summary

The PR enables vision for two Fireworks minimax-m3 cost-map entries and adds evidence-backed capability pins to the cost-map guard. The new missing-file check also rejects older cost-map branches that never contained the pins file.

Reviews (2) · Last reviewed commit: "fix(cost_map): fail the guard when the p..."

Comment thread ci_cd/cost_map_guard.py Outdated
@codspeed

codspeed Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing litellm_fireworks_minimax_m3_vision_pin (e5cc50e) with main (e73abe6)

Open in CodSpeed

@codecov

codecov Bot commented Sep 27, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@mateo-berri

Copy link
Copy Markdown
Contributor

@greptileai

Comment thread ci_cd/cost_map_guard.py
Comment on lines +135 to +136
if not pins_text:
return (f"{PINS_PATH} is missing; restore it, its pins were verified against live provider calls",)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Older branches fail the guard If a cost-map PR was cut before the pins file existed, its head lacks that file without having deleted it. This check rejects the PR even though merging it would retain the base branch’s pins.

def test_main_fails_a_pr_that_deletes_the_pins_file(tmp_path: Path) -> None:
subprocess.run(("git", "init", "-q", str(tmp_path)), check=True)
base: Final = _commit(tmp_path, BASE_MAP, "base", pins=PINS)
subprocess.run(("git", "rm", "-q", guard.PINS_PATH), cwd=tmp_path, check=True)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Unit test runs subprocesses This new deletion test runs git subprocesses, but tests/unit requires in-process tests with no subprocesses. That repository requirement must be satisfied before merging.

Context Used: AGENTS.md (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant