Retry transient Supabase errors in conversation status updates - #2192
Conversation
Transient Supabase infrastructure errors (5xx gateways, proxy DNS/cache overflows) while updating conversation status would leave a conversation stuck in a non-terminal state, logging 'Supabase error while updating conversation status' and 'Failed to mark conversation ... complete'. Wrap the RPC call in a tenacity retry (3 attempts, exponential backoff), mirroring the existing pattern in get_conversation_events. Mismatch (reassignment) errors are explicitly excluded from retry so the worker still exits cleanly when the row was reassigned. Signed-off-by: Claude <noreply@anthropic.com>
📂 Previous Runs
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Denied commands | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 26.4s | 4 | 8 | $0.2087 | 69,751 | 68,147 | 19,481 | 1,604 | 673 | 47,872 | 20,275 | 83 | — | — | src |
| ✅ | 101_loki_historical_logs_pod_deleted | 48.5s | 5 | 11 | $0.2741 | 91,933 | 88,840 | 21,455 | 3,093 | 972 | 66,489 | 22,351 | 535 | — | — | src |
| ✅ | 112_find_pvcs_by_uuid | 13.1s | 2 | 2 | $0.1559 | 33,167 | 32,257 | 17,935 | 910 | 473 | 14,319 | 17,938 | 280 | — | — | src |
| ✅ | 12_job_crashing | 28.5s | 4 | 9 | $0.2232 | 73,018 | 71,352 | 20,636 | 1,666 | 456 | 49,103 | 22,249 | 155 | — | — | src |
| ✅ | 176_network_policy_blocking_traffic_no_skills | 26.7s | 5 | 9 | $0.2123 | 86,578 | 85,044 | 19,141 | 1,534 | 520 | 65,446 | 19,598 | 231 | — | — | src |
| ✅ | 227_count_configmaps_per_namespace[0] | 13.2s | 3 | 6 | $0.1495 | 46,346 | 45,639 | 16,655 | 707 | 432 | 28,980 | 16,659 | 29 | — | — | src |
| ✅ | 243_pod_names_contain_service | 27.2s | 3 | 7 | $0.1884 | 49,780 | 48,168 | 18,141 | 1,612 | 580 | 29,362 | 18,806 | 251 | — | — | src |
| ✅ | 24_misconfigured_pvc | 27.2s | 5 | 10 | $0.2140 | 86,106 | 84,537 | 18,936 | 1,569 | 405 | 64,671 | 19,866 | 42 | — | — | src |
| ✅ | 254_elasticsearch_dr_test_log_check | 54.9s | 9 | 12 | $0.2733 | 118,571 | 115,049 | 17,555 | 3,522 | 794 | 96,746 | 18,303 | 202 | — | — | src |
| ✅ | 259_wrong_cluster_logs_confusion | 62.2s | 9 | 11 | $0.2903 | 124,993 | 121,081 | 17,909 | 3,912 | 1,450 | 102,602 | 18,479 | 779 | — | — | src |
| ✅ | 260_global_es_remote_cluster_logs | 66.3s | 10 | 13 | $0.3136 | 145,875 | 141,682 | 18,594 | 4,193 | 1,228 | 122,534 | 19,148 | 556 | — | — | src |
| ✅ | 43_current_datetime_from_prompt | 3.6s | 1 | — | $0.1017 | 14,428 | 14,306 | 14,306 | 122 | 122 | 0 | 14,306 | 78 | — | — | src |
| ✅ | 51_logs_summarize_errors | 17.4s | 3 | 2 | $0.1559 | 46,918 | 46,129 | 17,199 | 789 | 427 | 28,926 | 17,203 | 30 | — | — | src |
| ✅ | 61_exact_match_counting | 7.0s | 2 | 1 | $0.1146 | 29,149 | 28,928 | 14,642 | 221 | 152 | 14,283 | 14,645 | 34 | — | — | src |
| Total | 30.2s avg | 4.6 avg | 7.8 avg | $2.8754 | 1,016,613 | 991,159 | 21,455 | 25,454 | 1,450 | 731,333 | 259,826 | 3,285 | — | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 14 test/model combinations loaded
- master-27498497584 (created: 2026-06-14)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 17 test/model combinations loaded
- ci-benchmark-27491953079 (created: 2026-06-14)
Time comparison (seconds):
| Test case | This branch | master (2h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 26.4s | 25.9s | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 48.5s | 42.9s | ↑13% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 13.1s | 14.7s | ↓11% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 28.5s | 30.5s | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 26.7s | 26.0s | ±0% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 13.2s | 12.9s | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 27.2s | 25.4s | ±0% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 27.2s | 33.6s | ↓19% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 54.9s | 56.2s | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 62.2s | 57.0s | ±0% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 66.3s | 65.5s | ±0% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 3.6s | 3.0s | ↑22% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 17.4s | 16.2s | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 7.0s | 6.7s | ±0% | — | — |
| Total (all, n=14) | 30.2s | 29.7s | — | — | — |
| Comparable (m=14, b=0) | 30.2s | 29.7s | ±0% | — | — |
Cost comparison:
| Test case | This branch | master (2h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | $0.2087 | $0.2102 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | $0.2741 | $0.2428 | ↑13% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | $0.1559 | $0.1609 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | $0.2232 | $0.2304 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | $0.2123 | $0.2173 | ±0% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | $0.1495 | $0.1491 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | $0.1884 | $0.1878 | ±0% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | $0.2140 | $0.2384 | ↓10% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | $0.2733 | $0.2899 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | $0.2903 | $0.2518 | ↑15% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | $0.3136 | $0.3092 | ±0% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | $0.1017 | $0.0111 | ↑814% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | $0.1559 | $0.1556 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | $0.1146 | $0.1146 | ±0% | — | — |
| Total (all, n=14) | $0.2054 | $0.1978 | — | — | — |
| Comparable (m=14, b=0) | $0.2054 | $0.1978 | ±0% | — | — |
Total tokens comparison:
| Test case | This branch | master (2h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 69,751 | 69,863 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 91,933 | 71,262 | ↑29% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 33,167 | 33,544 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 73,018 | 73,006 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 86,578 | 70,749 | ↑22% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 46,346 | 46,342 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 49,780 | 49,682 | ±0% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 86,106 | 87,048 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 118,571 | 122,668 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 124,993 | 82,485 | ↑52% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 145,875 | 119,540 | ↑22% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 14,428 | 14,424 | ±0% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 46,918 | 46,897 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 29,149 | 29,149 | ±0% | — | — |
| Total (all, n=14) | 72,615 | 65,476 | — | — | — |
| Comparable (m=14, b=0) | 72,615 | 65,476 | ↑11% | — | — |
Cached tokens comparison:
| Test case | This branch | master (2h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 47,872 | 47,909 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 66,489 | 47,804 | ↑39% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 14,319 | 14,319 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 49,103 | 49,509 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 65,446 | 47,579 | ↑38% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 28,980 | 28,982 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 29,362 | 29,344 | ±0% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 64,671 | 63,701 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 96,746 | 99,549 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 102,602 | 61,908 | ↑66% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 122,534 | 95,513 | ↑28% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | 14,303 | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 28,926 | 28,927 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 14,283 | 14,283 | ±0% | — | — |
| Total (all, n=14) | 52,238 | 45,974 | — | — | — |
| Comparable (m=13, b=0) | 56,256 | 48,410 | ↑16% | — | — |
Turns comparison:
| Test case | This branch | master (2h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 4 | 4 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 5 | 4 | ↑25% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 4 | 4 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 5 | 4 | ↑25% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 3 | 3 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 3 | 3 | ±0% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 5 | 5 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 9 | 9 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 9 | 6 | ↑50% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 10 | 8 | ↑25% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 1 | 1 | ±0% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 3 | 3 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| Total (all, n=14) | 4.6 | 4.1 | — | — | — |
| Comparable (m=14, b=0) | 4.6 | 4.1 | ↑12% | — | — |
Tool calls comparison:
| Test case | This branch | master (2h ago) | Δ vs master | benchmark (6h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 8 | 8 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 11 | 9 | ↑22% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 9 | 11 | ↓18% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 9 | 9 | ±0% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 6 | 6 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 7 | 7 | ±0% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 10 | 11 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 12 | 13 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 11 | 8 | ↑38% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 13 | 13 | ±0% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 1 | 1 | ±0% | — | — |
| Total (all, n=14) | 7.2 | 7.7 | — | — | — |
| Comparable (m=13, b=0) | 7.8 | 7.7 | ±0% | — | — |
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/happy-euler-k1koo0 -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, fable-5, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/happy-euler-k1koo0 -f markers=regression -f filter=
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (2)
Walkthrough
ChangesRetry and error handling for update_conversation_status
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:0a4635122
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:0a4635122 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:0a4635122
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:0a4635122
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:0a4635122
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:0a4635122 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:0a4635122
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:0a4635122Patch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:0a4635122 \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:0a4635122Robusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:0a4635122 \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:0a4635122 |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
Wrap the result RPC in a tenacity @Retry (3 attempts, exponential 0.5-2s) that retries transient Supabase infra errors (DNS/cache overflows, 5xx) but NOT 'MISMATCH'/'not found' — those mean the row was reassigned/terminal, so drop calmly (return False, first result wins). Mirrors the update_conversation_status retry added in #2192. + 3 DAL-contract tests (retry-then-succeed, exhaust, no-retry-on-mismatch). Signed-off-by: Claude <noreply@anthropic.com>
Summary
Add automatic retry logic for transient infrastructure errors when updating conversation status in Supabase, while ensuring that MISMATCH errors (indicating the conversation was reassigned) are not retried and propagate immediately.
Changes
@retrydecorator that retries up to 3 times with exponential backoff (0.5s-2s) for transient errors like DNS failures, cache overflows, and 5xx gateway errorsretry_if_not_exception_type(ConversationReassignedError)to skip retries for MISMATCH errors, allowing the worker to exit cleanly when a conversation has been reassignedConversationReassignedErrorexceptions that bypass retry logic, while other exceptions are retried and eventually logged as failuresImplementation Details
_update()function decorated with@retryConversationReassignedErroris re-raised, other exceptions are logged and return Falsehttps://claude.ai/code/session_01L6SMoR2U3Rbg1spsjnxAAR
Summary by CodeRabbit
Bug Fixes