Skip to content

docs(blog): add July 21 progress updates to the Bedrock Invoke caching incident report - #628

Merged
mateo-berri merged 2 commits into
mainfrom
litellm_incident_post_progress_update
Jul 23, 2026
Merged

docs(blog): add July 21 progress updates to the Bedrock Invoke caching incident report#628
mateo-berri merged 2 commits into
mainfrom
litellm_incident_post_progress_update

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Updates the incident report with the progress shipped since it was published. The Known limitations item about Vertex AI and Azure is resolved: litellm#33807 (merged July 20) applies the same model-aware mid-conversation system handling as Bedrock Invoke on both paths, verified live, with e2e coverage. The weekly automated load test action is done: litellm#34166 (merged July 21) runs concurrent multi-turn sessions weekly against real Anthropic and Bedrock Invoke and fails on anomalies in error rate, warm-turn cache read share, cache writes, p95 turn latency, or recorded spend. The e2e action is marked partially done: the live-Bedrock cache regression tests from litellm#32963 and litellm#33807 are merged, while the full 250k-token scripted Claude Code session remains in progress and keeps its in-progress wording


Note

Low Risk
Documentation-only edits to an existing blog post; no runtime or API behavior changes.

Overview
Adds July 21 status notes to the Bedrock Invoke prompt-caching incident post so follow-up work is visible without rewriting the original commitments.

Under What we are changing, the e2e bullet now records that live cache-regression tests from #32963 (Bedrock) and #33807 (Vertex AI and Azure) are merged, while the full ~250k-token Claude Code session remains in progress. The weekly load-test bullet is marked done via #34166: CI runs concurrent Claude Code-shaped sessions weekly against real Anthropic and Bedrock Invoke and fails on baseline drift in errors, warm-turn cache reads, cache writes, p95 latency, or spend.

Under Known limitations, item 2 gets an update that Vertex AI and Azure needed the same model-aware mid-conversation system handling as Bedrock Invoke; #33807 ships that behavior with live verification and e2e coverage, resolving that open question.

Reviewed by Cursor Bugbot for commit f99cc48. Bugbot is set up for automated code reviews on this repo. Configure here.

@vercel

vercel Bot commented Jul 22, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
litellm Ready Ready Preview, Comment Jul 22, 2026 8:48pm

Request Review

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 0d9e0ce. Configure here.

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit f99cc48. Configure here.

@mateo-berri
mateo-berri merged commit d542cdc into main Jul 23, 2026
4 checks passed
@mateo-berri
mateo-berri deleted the litellm_incident_post_progress_update branch July 23, 2026 00:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant