Repository navigation
test(ci): refresh retired OpenAI tool-call models - #43676
Conversation
|
|
@greptileai Please review the latest commit, including the Bedrock routing assertion and its strengthened negative control |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
|
@greptileai Please review the current head. Both findings are addressed, with request validation and the original credential behavior preserved |
|
@greptileai Please review current head, including the verified worker-readiness repair and disambiguated response-format contract with retained negative controls |
|
bugbot run verbose=true Please review the current head for regressions in test isolation, request validation, and retained behavioral coverage |
|
@greptileai Please review the current head, including the worker response lifecycle and retained failure checks for missing spend records |
|
bugbot run verbose=true Please review the current head, including the worker response lifecycle and retained checks for missing spend records |
|
@greptileai Please review the current head, including held requests on both workers and the preserved exact-once survivor logging assertion |
|
bugbot run verbose=true Please review the current head, including held requests on both workers and the preserved exact-once survivor logging assertion |
|
@greptileai Please review the current head, including callback test isolation, the fixture package marker, and budget logger scoping with retained exception checks |
|
bugbot run verbose=true Please review the current head, including callback test isolation, the fixture package marker, and budget logger scoping with retained exception checks |
|
@greptileai Please review the current head, including resource ownership, nullable callback metadata, and preserved failure assertions in these test-only repairs |
|
bugbot run verbose=true Please review the current head for weakened assertions, resource cleanup, and accidental product changes in these test-only repairs |
|
@greptileai Please review the current head, including the budget permission fixture's supported decrease and unchanged authorization expectations across all roles |
|
bugbot run verbose=true Please review the current head, especially whether the budget fixture preserves authorization coverage without depending on case execution order |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 4d002f1. Configure here.
4d002f1 to
25b8ced
Compare
|
@greptileai Please review the current head after splitting this PR down to the live model updates |
TLDR
Problem this solves:
How it solves it:
This PR is now batch 1 only, limited to three test files. Legacy completion fakes, callback resets, schema snapshots, warning filters, and other fixture repairs have been removed from its diff
The streaming follow-up retains
temperature=0.2andseed=22, withreasoning_effort="none". A fresh real-provider call confirmed that combination is accepted. No live call is replaced with a fakeValidation
All four affected live cases pass with retries disabled.
make checkpasses the test-tree lint and test-quality budget gates. The replacement model exercises the Responses-to-chat transformation. Dropping returned tool calls in that active response transformation makes both added assertions fail; restoring the product passes all four cases. Hosted checks and reviews must qualify the new head independentlyThe two logging-file cases assert returned tool calls. These assertions do not establish delivery to the external logging services
Pre-Submission checklist
Screenshots / Proof of Fix
Use the isolated proxy at
http://127.0.0.1:49548. Registerci-tool-proofthroughPOST /model/new, settinglitellm_params.modelto the model under test and supplying the provider key privately. Both cases use the same weather-tool request below. Delete the temporary deployment afterwardBefore (9525452)
openai/gpt-3.5-turbo-1106and send the shared requestAfter (25b8ced)
openai/gpt-6-lunaand send the shared requestget_current_weathercalls, and 152 total tokensCaveats
Medium
Type
Test
Final Attestation