Skip to content

Fix eval-judge crash on truncated JSON (zeroed scores with no metrics) - #2187

Merged
aantn merged 5 commits into
masterfrom
claude/fix-eval-judge-truncated-json
Jun 14, 2026
Merged

aantn merged 5 commits into
masterfrom
claude/fix-eval-judge-truncated-json

Conversation

@aantn

@aantn aantn commented Jun 12, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

The eval correctness judge runs with use_cot, so its rationale and verdict share a single JSON tool call. autoevals' default max_tokens=512 truncates that JSON when the rationale is long (large evaluation outputs), and autoevals raises json.decoder.JSONDecodeError: Unterminated string with no error handling:

tests/llm/utils/property_manager.py:234: in update_test_results
    correctness_eval = evaluate_correctness(
tests/llm/utils/classifiers.py:229: in evaluate_correctness
    correctness_eval = classifier(
.venv/.../autoevals/llm.py:240: in _process_response
    args = json.loads(tool_call["function"]["arguments"])
E   json.decoder.JSONDecodeError: Unterminated string starting at: line 1 column 12 (char 11)

The crash happens before any score, cost, or token metric is recorded, so the test shows as failed with completely empty metric columns even when Holmes answered correctly. 95_skill_memory_leak_detection hit this in two consecutive CI runs on #2175 (its long memory-leak investigation produces a long judge rationale).

Fix

  • Raise the judge's max_tokens to 4096 so the rationale + verdict fit.
  • Retry the classifier call up to 3 times on JSONDecodeError for genuinely transient truncation.

Test-infra only; no product code touched.

https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo


Generated by Claude Code

Summary by CodeRabbit

  • Bug Fixes
    • Increased evaluation token capacity to prevent truncated judge responses and reduce JSON parsing failures.
    • Improved robustness of correctness scoring by treating missing outputs as empty strings, avoiding errors and reducing flakiness.

The correctness judge runs with use_cot, so its rationale and choice
share one JSON tool call. autoevals' default max_tokens=512 truncated
that JSON on long rationales (large evaluation outputs like
95_skill_memory_leak_detection, which crashed this way in two
consecutive CI runs), and autoevals raised json.JSONDecodeError before
any score or metric was recorded - the report showed the test failed
with empty cost/token columns even though Holmes answered correctly.

Raise the judge's max_tokens to 4096 and retry the classifier call up
to three times on malformed-JSON responses for genuinely transient
truncation.

https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo
Signed-off-by: Claude <noreply@anthropic.com>
aantn pushed a commit that referenced this pull request Jun 12, 2026
The grading fix moves to its own PR (#2187) so it can merge to master
independently of the SuggestSkills work. This reverts commit 4676d98.

https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

⚠️ 1 older run truncated

Older runs were omitted to stay under GitHub's 64KB comment size limit.


✅ Results of HolmesGPT evals

Automatically triggered by commit 083c4ad on branch claude/fix-eval-judge-truncated-json

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 14/14 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions Denied commands Src
✅ 09_crashpod 26.1s 4 8 $0.2119 70,388 68,767 19,526 1,621 664 48,029 20,738 94 — — src
✅ 101_loki_historical_logs_pod_deleted 37.8s 4 8 $0.2181 68,634 66,421 18,922 2,213 636 47,211 19,210 296 — — src
✅ 112_find_pvcs_by_uuid 14.2s 2 2 $0.1604 33,700 32,734 18,412 966 573 14,319 18,415 217 — — src
✅ 12_job_crashing 32.0s 5 10 $0.2395 92,768 90,863 20,683 1,905 513 68,903 21,960 169 — — src
✅ 176_network_policy_blocking_traffic_no_skills 27.4s 5 9 $0.2288 88,069 86,424 19,906 1,645 385 64,479 21,945 255 — — src
✅ 227_count_configmaps_per_namespace[0] 13.0s 3 6 $0.1491 46,341 45,630 16,645 711 437 28,981 16,649 30 — — src
✅ 243_pod_names_contain_service 25.2s 3 6 $0.1811 48,942 47,461 17,762 1,481 577 29,226 18,235 248 — — src
✅ 24_misconfigured_pvc 29.9s 5 11 $0.2300 88,342 86,461 19,431 1,881 747 65,452 21,009 122 — — src
✅ 254_elasticsearch_dr_test_log_check 58.2s 10 13 $0.2883 132,287 128,505 17,484 3,782 907 110,134 18,371 227 — — src
✅ 259_wrong_cluster_logs_confusion 48.5s 6 11 $0.2659 90,564 87,100 18,676 3,464 802 67,485 19,615 167 — — src
✅ 260_global_es_remote_cluster_logs 60.1s 10 14 $0.3057 145,385 141,748 18,610 3,637 708 121,118 20,630 193 — — src
✅ 43_current_datetime_from_prompt 3.4s 1 — $0.1017 14,428 14,306 14,306 122 122 0 14,306 78 — — src
✅ 51_logs_summarize_errors 16.7s 3 2 $0.1546 46,672 45,856 16,907 816 420 28,945 16,911 30 — — src
✅ 61_exact_match_counting 7.2s 2 1 $0.1146 29,147 28,926 14,640 221 152 14,283 14,643 34 — — src
Total 28.5s avg 4.5 avg 7.8 avg $2.8498 995,667 971,202 20,683 24,465 907 708,565 262,637 2,160 — —
Benchmark Comparison Details

Master baseline: latest master-* experiment (post-merge regression eval)
Status: 14 test/model combinations loaded

Benchmark baseline: latest ci-benchmark experiment on master
Status: 17 test/model combinations loaded

Time comparison (seconds):

Test case This branch master (1h ago) Δ vs master benchmark (1h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 26.1s 23.4s ↑12% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 37.8s 42.7s ↓12% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 14.2s 13.4s ±0% — —
12_job_crashing (opus-4.6) 📄 32.0s 30.7s ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 27.4s 25.9s ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 13.0s 15.2s ↓14% — —
243_pod_names_contain_service (opus-4.6) 📄 25.2s 29.3s ↓14% — —
24_misconfigured_pvc (opus-4.6) 📄 29.9s 28.9s ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 58.2s 58.1s ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 48.5s 49.4s ±0% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 60.1s 56.6s ±0% — —
43_current_datetime_from_prompt (opus-4.6) 📄 3.4s 3.0s ↑14% — —
51_logs_summarize_errors (opus-4.6) 📄 16.7s 16.4s ±0% — —
61_exact_match_counting (opus-4.6) 📄 7.2s 6.7s ±0% — —
Total (all, n=14) 28.5s 28.5s — — —
Comparable (m=14, b=0) 28.5s 28.5s ±0% — —

Cost comparison:

Test case This branch master (1h ago) Δ vs master benchmark (1h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 $0.2119 $0.2035 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 $0.2181 $0.2504 ↓13% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 $0.1604 $0.1482 ±0% — —
12_job_crashing (opus-4.6) 📄 $0.2395 $0.2466 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 $0.2288 $0.2132 ±0% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 $0.1491 $0.1486 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 $0.1811 $0.2022 ↓10% — —
24_misconfigured_pvc (opus-4.6) 📄 $0.2300 $0.2281 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 $0.2883 $0.2906 ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 $0.2659 $0.2283 ↑16% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 $0.3057 $0.2874 ±0% — —
43_current_datetime_from_prompt (opus-4.6) 📄 $0.1017 $0.1016 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 $0.1546 $0.1560 ±0% — —
61_exact_match_counting (opus-4.6) 📄 $0.1146 $0.1146 ±0% — —
Total (all, n=14) $0.2036 $0.2014 — — —
Comparable (m=14, b=0) $0.2036 $0.2014 ±0% — —

Total tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (1h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 70,388 68,236 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 68,634 88,763 ↓23% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 33,700 32,071 ±0% — —
12_job_crashing (opus-4.6) 📄 92,768 94,298 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 88,069 69,514 ↑27% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 46,341 46,338 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 48,942 67,520 ↓28% — —
24_misconfigured_pvc (opus-4.6) 📄 88,342 92,281 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 132,287 132,383 ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 90,564 67,686 ↑34% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 145,385 105,286 ↑38% — —
43_current_datetime_from_prompt (opus-4.6) 📄 14,428 14,424 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 46,672 47,004 ±0% — —
61_exact_match_counting (opus-4.6) 📄 29,147 29,148 ±0% — —
Total (all, n=14) 71,119 68,211 — — —
Comparable (m=14, b=0) 71,119 68,211 ±0% — —

Cached tokens comparison:

Test case This branch master (1h ago) Δ vs master benchmark (1h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 48,029 46,049 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 47,211 65,452 ↓28% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 14,319 14,319 ±0% — —
12_job_crashing (opus-4.6) 📄 68,903 69,492 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 64,479 46,932 ↑37% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 28,981 28,978 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 29,226 46,939 ↓38% — —
24_misconfigured_pvc (opus-4.6) 📄 65,452 69,120 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 110,134 110,172 ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 67,485 48,067 ↑40% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 121,118 82,021 ↑48% — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 28,945 28,922 ±0% — —
61_exact_match_counting (opus-4.6) 📄 14,283 14,283 ±0% — —
Total (all, n=14) 50,612 47,910 — — —
Comparable (m=13, b=0) 54,505 51,596 ±0% — —

Turns comparison:

Test case This branch master (1h ago) Δ vs master benchmark (1h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 4 4 ±0% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 4 5 ↓20% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 5 5 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 5 4 ↑25% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 3 3 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 3 4 ↓25% — —
24_misconfigured_pvc (opus-4.6) 📄 5 5 ±0% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 10 10 ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 6 5 ↑20% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 10 7 ↑43% — —
43_current_datetime_from_prompt (opus-4.6) 📄 1 1 ±0% — —
51_logs_summarize_errors (opus-4.6) 📄 3 3 ±0% — —
61_exact_match_counting (opus-4.6) 📄 2 2 ±0% — —
Total (all, n=14) 4.5 4.3 — — —
Comparable (m=14, b=0) 4.5 4.3 ±0% — —

Tool calls comparison:

Test case This branch master (1h ago) Δ vs master benchmark (1h ago) Δ vs benchmark
09_crashpod (opus-4.6) 📄 8 7 ↑14% — —
101_loki_historical_logs_pod_deleted (opus-4.6) 📄 8 9 ↓11% — —
112_find_pvcs_by_uuid (opus-4.6) 📄 2 2 ±0% — —
12_job_crashing (opus-4.6) 📄 10 11 ±0% — —
176_network_policy_blocking_traffic_no_skills (opus-4.6) 📄 9 10 ↓10% — —
227_count_configmaps_per_namespace[0] (opus-4.6) 📄 6 6 ±0% — —
243_pod_names_contain_service (opus-4.6) 📄 6 7 ↓14% — —
24_misconfigured_pvc (opus-4.6) 📄 11 10 ↑10% — —
254_elasticsearch_dr_test_log_check (opus-4.6) 📄 13 13 ±0% — —
259_wrong_cluster_logs_confusion (opus-4.6) 📄 11 6 ↑83% — —
260_global_es_remote_cluster_logs (opus-4.6) 📄 14 11 ↑27% — —
43_current_datetime_from_prompt (opus-4.6) 📄 — — — — —
51_logs_summarize_errors (opus-4.6) 📄 2 2 ±0% — —
61_exact_match_counting (opus-4.6) 📄 1 1 ±0% — —
Total (all, n=14) 7.2 7.3 — — —
Comparable (m=13, b=0) 7.8 7.3 ±0% — —

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-judge-truncated-json -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, multi-cluster, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, fable-5, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, gpt-5.5, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, opus-4.7, opus-4.8, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-eval-judge-truncated-json -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: f5ac8ef7-98ae-4104-866f-70ffa62cf820

📥 Commits

Reviewing files that changed from the base of the PR and between 3cbdb86 and e5382ef.

📒 Files selected for processing (1)
  • tests/llm/utils/classifiers.py

Walkthrough

Raises the correctness judge's LLMClassifier max_tokens to 16384 (with comments about avoiding truncation/JSONDecodeError) and normalizes output to "" by passing output or "" in both parent-span and non-parent-span classifier calls.

Changes

Correctness classifier tweaks

Layer / File(s) Summary
Increase correctness classifier max_tokens
tests/llm/utils/classifiers.py
Sets LLMClassifier max_tokens to 16384 and adds comments explaining this prevents truncation-related JSONDecodeError when the judge emits rationale/tool-call JSON.
Normalize output argument in both branches
tests/llm/utils/classifiers.py
Changes classifier invocations in the parent_span and non-parent_span paths to pass output or "" instead of output, ensuring None is converted to an empty string before scoring.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Suggested reviewers

  • moshemorad
  • Sheeproid
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main fix in the changeset: addressing eval-judge crashes caused by truncated JSON by increasing max_tokens to prevent truncation.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 100beb38b (built in 4m 45s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:100beb38b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:100beb38b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:100beb38b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:100beb38b
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:100beb38b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:100beb38b me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:100beb38b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:100beb38b

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:100beb38b \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:100beb38b

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:100beb38b \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:100beb38b

Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Jun 12, 2026

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 4b9370e
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a2bb83890daca00089c7ba4
😎 Deploy Preview https://deploy-preview-2187--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@netlify

netlify Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 083c4ad
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6a2e6394945d7000080cd7b0
😎 Deploy Preview https://deploy-preview-2187--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/llm/utils/classifiers.py`:
- Around line 257-259: _update the signature of _run_classifier_with_retry to
accept Optional[str] for the output parameter (matching evaluate_correctness)
and coerce None to an empty string before passing it to the classifier;
specifically, change the type hint on _run_classifier_with_retry (currently def
_run_classifier_with_retry(classifier, prompt_prefix: str, output: str,
expected_elements_str: str)) to accept Optional[str], then inside the function
set a local variable (e.g., safe_output = output or "") and use that safe_output
when invoking classifier and anywhere else output is read; keep all other logic
intact.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 409c8727-df20-470c-8065-20c0981bd227

📥 Commits

Reviewing files that changed from the base of the PR and between 30a5e2b and 254fbcb.

📒 Files selected for processing (1)
  • tests/llm/utils/classifiers.py

Comment thread tests/llm/utils/classifiers.py Outdated
evaluate_correctness takes output: Optional[str] and passed it straight
through; coerce None to an empty string before invoking the classifier
(CodeRabbit review finding).

Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) June 12, 2026 19:30
claude and others added 2 commits June 12, 2026 19:31
The truncated-JSON crash is fully explained by autoevals' default
max_tokens=512 cutting off the judge's chain-of-thought tool call, so
raising the limit is the whole fix. Bump it from 4096 to 16384 — far
above any realistic rationale length — and remove the JSONDecodeError
retry helper, which added complexity without addressing a distinct
failure mode.

https://claude.ai/code/session_01VLkU3Rj1FoXXpkmNAyS6vo
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn merged commit 4ff71e4 into master Jun 14, 2026
18 of 19 checks passed
@aantn
aantn deleted the claude/fix-eval-judge-truncated-json branch June 14, 2026 08:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants