Skip to content

test(ci): refresh retired OpenAI tool-call models - #43676

Merged
yuneng-berri merged 1 commit into
mainfrom
litellm_cci_stale_tests_20260929
Sep 30, 2026
Merged

yuneng-berri merged 1 commit into
mainfrom
litellm_cci_stale_tests_20260929

Conversation

@yuneng-berri

@yuneng-berri yuneng-berri commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

Problem this solves:

  • Four live tool-call cases use a retired OpenAI model

How it solves it:

  • Select a supported model and retain live provider calls
  • Assert returned tool calls are present and correctly named
  • Preserve sampling parameters with explicit non-reasoning mode

This PR is now batch 1 only, limited to three test files. Legacy completion fakes, callback resets, schema snapshots, warning filters, and other fixture repairs have been removed from its diff

The streaming follow-up retains temperature=0.2 and seed=22, with reasoning_effort="none". A fresh real-provider call confirmed that combination is accepted. No live call is replaced with a fake

Validation

All four affected live cases pass with retries disabled. make check passes the test-tree lint and test-quality budget gates. The replacement model exercises the Responses-to-chat transformation. Dropping returned tool calls in that active response transformation makes both added assertions fail; restoring the product passes all four cases. Hosted checks and reviews must qualify the new head independently

The two logging-file cases assert returned tool calls. These assertions do not establish delivery to the external logging services

Pre-Submission checklist

  • I have added meaningful tests
  • The handful of test files covering my change pass locally
  • My PR passes all required CI/CD checks
  • My PR's scope is as isolated as possible
  • I have received the required current-head Greptile confidence score

Screenshots / Proof of Fix

Use the isolated proxy at http://127.0.0.1:49548. Register ci-tool-proof through POST /model/new, setting litellm_params.model to the model under test and supplying the provider key privately. Both cases use the same weather-tool request below. Delete the temporary deployment afterward

curl -sS http://127.0.0.1:49548/v1/chat/completions \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H 'Content-Type: application/json' \
  --data-binary @- <<'JSON'
{"model":"ci-tool-proof","messages":[{"role":"user","content":"What's the weather like in San Francisco, Tokyo, and Paris?"}],"tools":[{"type":"function","function":{"name":"get_current_weather","description":"Get the current weather in a given location","parameters":{"type":"object","properties":{"location":{"type":"string"},"unit":{"type":"string","enum":["celsius","fahrenheit"]}},"required":["location"]}}}],"tool_choice":"auto"}
JSON

Before (9525452)

  1. Register openai/gpt-3.5-turbo-1106 and send the shared request
  2. Observe HTTP 404 and the provider message that the model is deprecated

After (25b8ced)

  1. Register openai/gpt-6-luna and send the shared request
  2. Observe HTTP 200, three get_current_weather calls, and 152 total tokens

Caveats

Medium

  • CircleCI pipeline creation is currently disabled for this project
  • OSV reports advisories in unchanged dependency lockfiles
  • The unchanged scheduler stagger test fails in hosted CI
  • The miscellaneous unit-test job also fails in hosted CI

Type

Test

Final Attestation

  • The tests check the right things, including edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@yuneng-berri
yuneng-berri requested a review from a team September 29, 2026 08:20
@greptile-apps

greptile-apps Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Low risk] Test files update model names and add assertions.

The current three-file PR appears safe to merge based on this review.

Summary

The PR replaces a retired OpenAI model in four live tool-calling cases, adds tool-call assertions to two tests, and preserves sampling parameters in the streaming follow-up. All previous Greptile threads are resolved and concern changes outside the current three-file PR.

Reviews (10) · Last reviewed commit: "test(ci): refresh retired OpenAI tool-ca..."

Comment thread proxy_server_config.yaml Outdated
Comment thread tests/unit/interactions/fixtures/gemini_interactions_contract.json Outdated
@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the latest commit, including the Bedrock routing assertion and its strengthened negative control

@codecov

codecov Bot commented Sep 29, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current head. Both findings are addressed, with request validation and the original credential behavior preserved

Comment thread tests/unit/interactions/fixtures/gemini_interactions_contract.json Outdated
@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review current head, including the verified worker-readiness repair and disambiguated response-format contract with retained negative controls

Comment thread tests/integration/observability/test_passthrough_upstream_error_chaos.py Outdated
@yuneng-berri

Copy link
Copy Markdown
Contributor Author

bugbot run verbose=true

Please review the current head for regressions in test isolation, request validation, and retained behavioral coverage

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread tests/integration/observability/test_passthrough_upstream_error_chaos.py Outdated
@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current head, including the worker response lifecycle and retained failure checks for missing spend records

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

bugbot run verbose=true

Please review the current head, including the worker response lifecycle and retained checks for missing spend records

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread tests/integration/observability/test_passthrough_upstream_error_chaos.py Outdated
@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current head, including held requests on both workers and the preserved exact-once survivor logging assertion

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

bugbot run verbose=true

Please review the current head, including held requests on both workers and the preserved exact-once survivor logging assertion

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current head, including callback test isolation, the fixture package marker, and budget logger scoping with retained exception checks

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

bugbot run verbose=true

Please review the current head, including callback test isolation, the fixture package marker, and budget logger scoping with retained exception checks

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current head, including resource ownership, nullable callback metadata, and preserved failure assertions in these test-only repairs

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

bugbot run verbose=true

Please review the current head for weakened assertions, resource cleanup, and accidental product changes in these test-only repairs

@cursor cursor Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current head, including the budget permission fixture's supported decrease and unchanged authorization expectations across all roles

@yuneng-berri

Copy link
Copy Markdown
Contributor Author

bugbot run verbose=true

Please review the current head, especially whether the budget fixture preserves authorization coverage without depending on case execution order

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 4d002f1. Configure here.

@yuneng-berri
yuneng-berri force-pushed the litellm_cci_stale_tests_20260929 branch from 4d002f1 to 25b8ced Compare September 30, 2026 00:47
@yuneng-berri yuneng-berri changed the title test(ci): repair stale fixtures in scheduled CircleCI suites test(ci): refresh retired OpenAI tool-call models Sep 30, 2026
@yuneng-berri

Copy link
Copy Markdown
Contributor Author

@greptileai Please review the current head after splitting this PR down to the live model updates

@yuneng-berri
yuneng-berri merged commit 7d9cc28 into main Sep 30, 2026
100 of 103 checks passed
@yuneng-berri
yuneng-berri deleted the litellm_cci_stale_tests_20260929 branch September 30, 2026 04:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants