Fix eval-judge crash on truncated JSON (zeroed scores with no metrics) - #2187
Conversation
The correctness judge runs with use_cot, so its rationale and choice share one JSON tool call. autoevals' default max_tokens=512 truncated that JSON on long rationales (large evaluation outputs like 95_skill_memory_leak_detection, which crashed this way in two consecutive CI runs), and autoevals raised json.JSONDecodeError before any score or metric was recorded - the report showed the test failed with empty cost/token columns even though Holmes answered correctly. Raise the judge's max_tokens to 4096 and retry the classifier call up to three times on malformed-JSON responses for genuinely transient truncation. https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo Signed-off-by: Claude <noreply@anthropic.com>
The grading fix moves to its own PR (#2187) so it can merge to master independently of the SuggestSkills work. This reverts commit 4676d98. https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo Signed-off-by: Claude <noreply@anthropic.com>
📂 Previous Runs
|
| Status | Test case | Time | Turns | Tools | Cost | Total tokens | Input | Max input | Output | Max output | Cached | Non-cached | Reasoning | Compactions | Denied commands | Src |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ✅ | 09_crashpod | 26.1s | 4 | 8 | $0.2119 | 70,388 | 68,767 | 19,526 | 1,621 | 664 | 48,029 | 20,738 | 94 | — | — | src |
| ✅ | 101_loki_historical_logs_pod_deleted | 37.8s | 4 | 8 | $0.2181 | 68,634 | 66,421 | 18,922 | 2,213 | 636 | 47,211 | 19,210 | 296 | — | — | src |
| ✅ | 112_find_pvcs_by_uuid | 14.2s | 2 | 2 | $0.1604 | 33,700 | 32,734 | 18,412 | 966 | 573 | 14,319 | 18,415 | 217 | — | — | src |
| ✅ | 12_job_crashing | 32.0s | 5 | 10 | $0.2395 | 92,768 | 90,863 | 20,683 | 1,905 | 513 | 68,903 | 21,960 | 169 | — | — | src |
| ✅ | 176_network_policy_blocking_traffic_no_skills | 27.4s | 5 | 9 | $0.2288 | 88,069 | 86,424 | 19,906 | 1,645 | 385 | 64,479 | 21,945 | 255 | — | — | src |
| ✅ | 227_count_configmaps_per_namespace[0] | 13.0s | 3 | 6 | $0.1491 | 46,341 | 45,630 | 16,645 | 711 | 437 | 28,981 | 16,649 | 30 | — | — | src |
| ✅ | 243_pod_names_contain_service | 25.2s | 3 | 6 | $0.1811 | 48,942 | 47,461 | 17,762 | 1,481 | 577 | 29,226 | 18,235 | 248 | — | — | src |
| ✅ | 24_misconfigured_pvc | 29.9s | 5 | 11 | $0.2300 | 88,342 | 86,461 | 19,431 | 1,881 | 747 | 65,452 | 21,009 | 122 | — | — | src |
| ✅ | 254_elasticsearch_dr_test_log_check | 58.2s | 10 | 13 | $0.2883 | 132,287 | 128,505 | 17,484 | 3,782 | 907 | 110,134 | 18,371 | 227 | — | — | src |
| ✅ | 259_wrong_cluster_logs_confusion | 48.5s | 6 | 11 | $0.2659 | 90,564 | 87,100 | 18,676 | 3,464 | 802 | 67,485 | 19,615 | 167 | — | — | src |
| ✅ | 260_global_es_remote_cluster_logs | 60.1s | 10 | 14 | $0.3057 | 145,385 | 141,748 | 18,610 | 3,637 | 708 | 121,118 | 20,630 | 193 | — | — | src |
| ✅ | 43_current_datetime_from_prompt | 3.4s | 1 | — | $0.1017 | 14,428 | 14,306 | 14,306 | 122 | 122 | 0 | 14,306 | 78 | — | — | src |
| ✅ | 51_logs_summarize_errors | 16.7s | 3 | 2 | $0.1546 | 46,672 | 45,856 | 16,907 | 816 | 420 | 28,945 | 16,911 | 30 | — | — | src |
| ✅ | 61_exact_match_counting | 7.2s | 2 | 1 | $0.1146 | 29,147 | 28,926 | 14,640 | 221 | 152 | 14,283 | 14,643 | 34 | — | — | src |
| Total | 28.5s avg | 4.5 avg | 7.8 avg | $2.8498 | 995,667 | 971,202 | 20,683 | 24,465 | 907 | 708,565 | 262,637 | 2,160 | — | — |
Benchmark Comparison Details
Master baseline: latest master-* experiment (post-merge regression eval)
Status: 14 test/model combinations loaded
- master-27491196671 (created: 2026-06-14)
Benchmark baseline: latest ci-benchmark experiment on master
Status: 17 test/model combinations loaded
- ci-benchmark-27491953079 (created: 2026-06-14)
Time comparison (seconds):
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (1h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 26.1s | 23.4s | ↑12% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 37.8s | 42.7s | ↓12% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 14.2s | 13.4s | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 32.0s | 30.7s | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 27.4s | 25.9s | ±0% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 13.0s | 15.2s | ↓14% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 25.2s | 29.3s | ↓14% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 29.9s | 28.9s | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 58.2s | 58.1s | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 48.5s | 49.4s | ±0% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 60.1s | 56.6s | ±0% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 3.4s | 3.0s | ↑14% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 16.7s | 16.4s | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 7.2s | 6.7s | ±0% | — | — |
| Total (all, n=14) | 28.5s | 28.5s | — | — | — |
| Comparable (m=14, b=0) | 28.5s | 28.5s | ±0% | — | — |
Cost comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (1h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | $0.2119 | $0.2035 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | $0.2181 | $0.2504 | ↓13% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | $0.1604 | $0.1482 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | $0.2395 | $0.2466 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | $0.2288 | $0.2132 | ±0% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | $0.1491 | $0.1486 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | $0.1811 | $0.2022 | ↓10% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | $0.2300 | $0.2281 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | $0.2883 | $0.2906 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | $0.2659 | $0.2283 | ↑16% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | $0.3057 | $0.2874 | ±0% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | $0.1017 | $0.1016 | ±0% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | $0.1546 | $0.1560 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | $0.1146 | $0.1146 | ±0% | — | — |
| Total (all, n=14) | $0.2036 | $0.2014 | — | — | — |
| Comparable (m=14, b=0) | $0.2036 | $0.2014 | ±0% | — | — |
Total tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (1h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 70,388 | 68,236 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 68,634 | 88,763 | ↓23% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 33,700 | 32,071 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 92,768 | 94,298 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 88,069 | 69,514 | ↑27% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 46,341 | 46,338 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 48,942 | 67,520 | ↓28% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 88,342 | 92,281 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 132,287 | 132,383 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 90,564 | 67,686 | ↑34% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 145,385 | 105,286 | ↑38% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 14,428 | 14,424 | ±0% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 46,672 | 47,004 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 29,147 | 29,148 | ±0% | — | — |
| Total (all, n=14) | 71,119 | 68,211 | — | — | — |
| Comparable (m=14, b=0) | 71,119 | 68,211 | ±0% | — | — |
Cached tokens comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (1h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 48,029 | 46,049 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 47,211 | 65,452 | ↓28% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 14,319 | 14,319 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 68,903 | 69,492 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 64,479 | 46,932 | ↑37% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 28,981 | 28,978 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 29,226 | 46,939 | ↓38% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 65,452 | 69,120 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 110,134 | 110,172 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 67,485 | 48,067 | ↑40% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 121,118 | 82,021 | ↑48% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 28,945 | 28,922 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 14,283 | 14,283 | ±0% | — | — |
| Total (all, n=14) | 50,612 | 47,910 | — | — | — |
| Comparable (m=13, b=0) | 54,505 | 51,596 | ±0% | — | — |
Turns comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (1h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 4 | 4 | ±0% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 4 | 5 | ↓20% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 5 | 5 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 5 | 4 | ↑25% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 3 | 3 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 3 | 4 | ↓25% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 5 | 5 | ±0% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 10 | 10 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 6 | 5 | ↑20% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 10 | 7 | ↑43% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | 1 | 1 | ±0% | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 3 | 3 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| Total (all, n=14) | 4.5 | 4.3 | — | — | — |
| Comparable (m=14, b=0) | 4.5 | 4.3 | ±0% | — | — |
Tool calls comparison:
| Test case | This branch | master (1h ago) | Δ vs master | benchmark (1h ago) | Δ vs benchmark |
|---|---|---|---|---|---|
| 09_crashpod (opus-4.6) 📄 | 8 | 7 | ↑14% | — | — |
| 101_loki_historical_logs_pod_deleted (opus-4.6) 📄 | 8 | 9 | ↓11% | — | — |
| 112_find_pvcs_by_uuid (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| 12_job_crashing (opus-4.6) 📄 | 10 | 11 | ±0% | — | — |
| 176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 | 9 | 10 | ↓10% | — | — |
| 227_count_configmaps_per_namespace[0] (opus-4.6) 📄 | 6 | 6 | ±0% | — | — |
| 243_pod_names_contain_service (opus-4.6) 📄 | 6 | 7 | ↓14% | — | — |
| 24_misconfigured_pvc (opus-4.6) 📄 | 11 | 10 | ↑10% | — | — |
| 254_elasticsearch_dr_test_log_check (opus-4.6) 📄 | 13 | 13 | ±0% | — | — |
| 259_wrong_cluster_logs_confusion (opus-4.6) 📄 | 11 | 6 | ↑83% | — | — |
| 260_global_es_remote_cluster_logs (opus-4.6) 📄 | 14 | 11 | ↑27% | — | — |
| 43_current_datetime_from_prompt (opus-4.6) 📄 | — | — | — | — | — |
| 51_logs_summarize_errors (opus-4.6) 📄 | 2 | 2 | ±0% | — | — |
| 61_exact_match_counting (opus-4.6) 📄 | 1 | 1 | ±0% | — | — |
| Total (all, n=14) | 7.2 | 7.3 | — | — | — |
| Comparable (m=13, b=0) | 7.8 | 7.3 | ±0% | — | — |
Comparison indicators:
±0%— diff under 10% (within noise threshold)↑N%/↓N%— diff 10-25%↑N%/↓N%— diff over 25% (significant)
📖 Legend
| Icon | Meaning |
|---|---|
| ✅ | The test was successful |
| ➖ | The test was skipped |
| The test failed but is known to be flaky or known to fail | |
| 🚧 | The test had a setup failure (not a code regression) |
| 🔧 | The test failed due to mock data issues (not a code regression) |
| 🚫 | The test was throttled by API rate limits/overload |
| ❌ | The test failed and should be fixed before merging the PR |
🔄 Re-run evals manually
⚠️ Warning:/evalcomments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.To test workflow changes, use the GitHub CLI or Actions UI instead:
gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-judge-truncated-json -f markers=regression -f filter=
Option 1: Comment on this PR with /eval:
/eval
tags: regression
Or with more options (one per line):
/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5
Run evals on a different branch (e.g., master) for comparison:
/eval
branch: master
tags: regression
| Option | Description |
|---|---|
model |
Model(s) to test (default: same as automatic runs) |
tags |
Pytest tags / markers (no default - runs all tests!) |
id |
Eval ID / pytest -k filter (use /list to see valid eval names) |
iterations |
Number of runs, max 10 |
branch |
Run evals on a different branch (for cross-branch comparison) |
Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.
Option 2: Trigger via GitHub Actions UI → "Run workflow"
Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):
| Label | Effect |
|---|---|
evals-tag-<name> |
Run tests with tag <name> alongside regression |
evals-id-<name> |
Run a specific eval by test ID |
evals-model-<name> |
Override the model (use model list name, e.g. sonnet-4.5) |
Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5
🏷️ Valid tags
benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs
🤖 Valid models
deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, fable-5, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6
Commands: /eval · /rerun · /list
CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-judge-truncated-json -f markers=regression -f filter=
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
WalkthroughRaises the correctness judge's ChangesCorrectness classifier tweaks
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Suggested reviewers
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:100beb38b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:100beb38b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:100beb38b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:100beb38b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:100beb38b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:100beb38b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:100beb38b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:100beb38bPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:100beb38b \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:100beb38bRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:100beb38b \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:100beb38b |
Signed-off-by: Claude <noreply@anthropic.com>
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/llm/utils/classifiers.py`:
- Around line 257-259: _update the signature of _run_classifier_with_retry to
accept Optional[str] for the output parameter (matching evaluate_correctness)
and coerce None to an empty string before passing it to the classifier;
specifically, change the type hint on _run_classifier_with_retry (currently def
_run_classifier_with_retry(classifier, prompt_prefix: str, output: str,
expected_elements_str: str)) to accept Optional[str], then inside the function
set a local variable (e.g., safe_output = output or "") and use that safe_output
when invoking classifier and anywhere else output is read; keep all other logic
intact.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 409c8727-df20-470c-8065-20c0981bd227
📒 Files selected for processing (1)
tests/llm/utils/classifiers.py
evaluate_correctness takes output: Optional[str] and passed it straight through; coerce None to an empty string before invoking the classifier (CodeRabbit review finding). Signed-off-by: Claude <noreply@anthropic.com>
The truncated-JSON crash is fully explained by autoevals' default max_tokens=512 cutting off the judge's chain-of-thought tool call, so raising the limit is the whole fix. Bump it from 4096 to 16384 — far above any realistic rationale length — and remove the JSONDecodeError retry helper, which added complexity without addressing a distinct failure mode. https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo Signed-off-by: Claude <noreply@anthropic.com>
Problem
The eval correctness judge runs with
use_cot, so its rationale and verdict share a single JSON tool call. autoevals' defaultmax_tokens=512truncates that JSON when the rationale is long (large evaluation outputs), and autoevals raisesjson.decoder.JSONDecodeError: Unterminated stringwith no error handling:The crash happens before any score, cost, or token metric is recorded, so the test shows as failed with completely empty metric columns even when Holmes answered correctly.
95_skill_memory_leak_detectionhit this in two consecutive CI runs on #2175 (its long memory-leak investigation produces a long judge rationale).Fix
max_tokensto 4096 so the rationale + verdict fit.JSONDecodeErrorfor genuinely transient truncation.Test-infra only; no product code touched.
https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo
Generated by Claude Code
Summary by CodeRabbit