Skip to content

Enable strict tool calling universally with per-tool compatibility checks - #1790

Merged
aantn merged 7 commits into
masterfrom
claude/tdd-mcp-json-schemas-4DkIW
Mar 16, 2026
Merged

aantn merged 7 commits into
masterfrom
claude/tdd-mcp-json-schemas-4DkIW

Conversation

@aantn

@aantn aantn commented Mar 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR refactors strict tool calling from a model-based allowlist to a universal default with automatic per-tool compatibility detection. Strict mode is now enabled by default for all models and can only be disabled globally via environment variable. Tools with dynamic-key parameters are automatically excluded from strict mode on a per-tool basis.

Key Changes

  • Simplified strict mode configuration: Replaced LLMS_WITH_STRICT_TOOL_CALLS model allowlist with HOLMES_DISABLE_STRICT_TOOL_CALLS boolean flag. Strict mode is now enabled universally by default.

  • Added is_strict_compatible() method to ToolParameter: Recursively checks if a parameter and all nested parameters can be used in strict mode. Parameters with dynamic keys (additionalProperties set to a schema or True) are marked as incompatible.

  • Implemented _is_tool_strict_compatible() helper: Validates all parameters in a tool before enabling strict mode. Tools with any incompatible parameters are automatically excluded from strict mode.

  • Improved additionalProperties handling in schema generation:

    • Objects with explicit properties now set additionalProperties: false only in strict mode
    • Objects with dynamic-key schemas preserve the additionalProperties schema definition
    • Prevents strict mode from being applied to tools that require dynamic keys
  • Added additional_properties field to ToolParameter: Stores the additionalProperties JSON Schema value (None, False, or a schema dict) for proper schema preservation and strict mode compatibility checking.

  • Updated MCP toolset parser: Now extracts and preserves additionalProperties from MCP tool schemas, resolving any nested $ref or anyOf references.

  • Updated documentation: Clarified that strict mode is now universal with automatic per-tool exclusions for dynamic-key parameters.

Implementation Details

The strict mode logic now follows this flow:

  1. Check if STRICT_TOOL_CALLS_ENABLED is true (default unless HOLMES_DISABLE_STRICT_TOOL_CALLS=true)
  2. For each tool, verify all parameters are strict-compatible via _is_tool_strict_compatible()
  3. Only enable strict mode if both conditions are met
  4. When generating schemas in strict mode, set additionalProperties: false only on objects with explicit properties, preserving dynamic-key schemas where needed

This ensures OpenAI and Anthropic's strict mode requirements (all objects must have additionalProperties: false) are met while still supporting tools with dynamic parameters by excluding them from strict mode.

https://claude.ai/code/session_01UJPFbZ7Y33QNGRHFDP4sRr

Summary by CodeRabbit

  • New Features

    • New env var TOOL_SCHEMA_NO_PARAM_OBJECT_IF_NO_PARAMS (default false) to omit empty parameter objects.
    • Automatic coercion of tool-call parameters to match declared JSON Schema types (e.g., parsing stringified JSON, wrapping single values into arrays).
  • Bug Fixes

    • Strict tool-calling is enabled by default; disable with HOLMES_DISABLE_STRICT_TOOL_CALLS.
    • Strict-mode now respects per-tool exceptions for dynamic-key or nested parameter schemas.
    • Tool OpenAI-format generation no longer depends on a target model.
  • Documentation

    • Env var docs updated; LLMS_WITH_STRICT_TOOL_CALLS retired and replaced by HOLMES_DISABLE_STRICT_TOOL_CALLS.
  • Tests

    • Added extensive tests for coercion and strict-mode behavior.

… dynamic keys

Replace LLMS_WITH_STRICT_TOOL_CALLS (model-based opt-in) with universal strict mode
enabled by default. Tools with additionalProperties schemas (dynamic-key objects like
filter maps) are automatically excluded from strict mode per-tool, since both OpenAI
and Anthropic require additionalProperties: false on all objects in strict mode.

Changes:
- Add additional_properties field and is_strict_compatible() to ToolParameter
- openai_formatting: enable strict mode for all models, auto-disable per-tool
  when additionalProperties has a schema value
- Preserve additionalProperties in MCP _parse_tool_parameter
- Replace LLMS_WITH_STRICT_TOOL_CALLS env var with HOLMES_DISABLE_STRICT_TOOL_CALLS
- Add unit tests for strict compatibility checks and per-tool toggling

https://claude.ai/code/session_01UJPFbZ7Y33QNGRHFDP4sRr
Signed-off-by: Claude <noreply@anthropic.com>
@claude

claude Bot commented Mar 15, 2026

Copy link
Copy Markdown
Contributor

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review.

@github-actions

github-actions Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ 180a6de (#23150504820)

✅ Results of HolmesGPT evals

Automatically triggered by commit 180a6de on branch claude/tdd-mcp-json-schemas-4DkIW (labels: evals-tag-mcp)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 41.3s 8 12 $0.2933 167,362 164,941 24,732 2,421 665 138,404 26,537 — —
✅ 101_loki_historical_logs_pod_deleted 41.0s 6 11 $0.2699 128,873 126,235 23,834 2,638 861 101,560 24,675 — —
✅ 111_pod_names_contain_service 35.0s 7 10 $0.2567 141,552 139,512 22,916 2,040 579 115,674 23,838 — —
✅ 112_find_pvcs_by_uuid 23.7s 4 6 $0.2125 82,992 81,499 23,109 1,493 704 58,048 23,451 — —
✅ 12_job_crashing 29.5s 5 11 $0.2369 107,109 105,204 23,556 1,905 585 81,374 23,830 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 39.9s 6 14 $0.4420 135,742 133,155 26,723 2,587 619 78,757 54,398 — —
✅ 227_count_configmaps_per_namespace[0] 25.5s 6 10 $0.3182 113,627 112,112 20,654 1,515 519 73,000 39,112 — —
✅ 24_misconfigured_pvc 35.1s 6 13 $0.2581 123,898 121,618 23,534 2,280 706 96,976 24,642 — —
✅ 43_current_datetime_from_prompt 4.5s 1 — $0.1088 17,063 16,945 16,945 118 118 0 16,945 — —
✅ 61_exact_match_counting 14.4s 4 4 $0.1522 71,291 70,820 18,252 471 237 52,556 18,264 — —
Total 29.0s avg 5.3 avg 10.1 avg $2.5486 1,089,509 1,072,041 26,723 17,468 861 796,349 275,692 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 68b14da (#23147165345)

✅ Results of HolmesGPT evals

Automatically triggered by commit 68b14da on branch claude/tdd-mcp-json-schemas-4DkIW (labels: evals-tag-mcp)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.3s 5 10 $0.2451 105,773 103,723 23,781 2,050 791 78,776 24,947 — —
✅ 101_loki_historical_logs_pod_deleted 53.0s 7 15 $0.4838 169,885 166,515 31,433 3,370 669 110,887 55,628 — —
✅ 111_pod_names_contain_service 30.2s 5 9 $0.2327 103,157 101,163 22,775 1,994 608 78,095 23,068 — —
✅ 112_find_pvcs_by_uuid 23.5s 5 5 $0.2185 101,001 99,611 23,024 1,390 397 76,245 23,366 — —
✅ 12_job_crashing 32.9s 5 15 $0.2759 116,652 114,242 26,176 2,410 869 86,275 27,967 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 35.7s 6 12 $0.2838 133,400 130,785 26,008 2,615 897 104,004 26,781 — —
✅ 227_count_configmaps_per_namespace[0] 24.5s 6 11 $0.2180 117,124 115,593 21,195 1,531 519 94,384 21,209 — —
✅ 24_misconfigured_pvc 35.5s 6 15 $0.2667 126,809 124,362 24,112 2,447 737 99,196 25,166 — —
✅ 43_current_datetime_from_prompt 4.2s 1 — $0.1112 17,450 17,334 17,334 116 116 0 17,334 — —
✅ 61_exact_match_counting 10.2s 3 2 $0.1389 53,473 53,133 18,046 340 193 35,076 18,057 — —
Total 28.0s avg 4.9 avg 10.4 avg $2.4747 1,044,724 1,026,461 31,433 18,263 897 762,938 263,523 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 68b14da (#23144789869)

✅ Results of HolmesGPT evals

Automatically triggered by commit 68b14da on branch claude/tdd-mcp-json-schemas-4DkIW

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 27.1s 4 8 $0.2193 82,821 80,995 22,711 1,826 708 57,705 23,290 — —
✅ 101_loki_historical_logs_pod_deleted 36.0s 5 9 $0.2434 106,216 103,933 22,922 2,283 872 80,451 23,482 — —
✅ 111_pod_names_contain_service 27.5s 4 8 $0.2132 81,523 79,791 22,161 1,732 668 57,059 22,732 — —
✅ 112_find_pvcs_by_uuid 25.1s 5 6 $0.2260 105,878 104,410 23,527 1,468 492 80,484 23,926 — —
✅ 12_job_crashing 38.7s 6 16 $0.3012 143,552 140,800 27,288 2,752 748 112,387 28,413 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 36.3s 6 12 $0.2696 128,842 126,437 24,858 2,405 803 100,837 25,600 — —
✅ 227_count_configmaps_per_namespace[0] 22.8s 5 10 $0.2053 97,557 96,110 20,895 1,447 519 74,999 21,111 — —
✅ 24_misconfigured_pvc 35.9s 6 15 $0.2705 127,522 125,051 24,284 2,471 722 99,312 25,739 — —
✅ 43_current_datetime_from_prompt 4.5s 1 — $0.1112 17,450 17,334 17,334 116 116 0 17,334 — —
✅ 61_exact_match_counting 14.9s 4 4 $0.1554 72,877 72,400 18,652 477 239 53,736 18,664 — —
Total 26.9s avg 4.6 avg 9.8 avg $2.2151 964,238 947,261 27,288 16,977 872 716,970 230,291 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 38cbc93 (#23143224753)

✅ Results of HolmesGPT evals

Automatically triggered by commit 38cbc93 on branch claude/tdd-mcp-json-schemas-4DkIW

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.4s 5 10 $0.2408 106,396 104,371 23,764 2,025 574 80,311 24,060 — —
✅ 101_loki_historical_logs_pod_deleted 54.5s 8 14 $0.4771 190,917 187,658 29,462 3,259 610 134,756 52,902 — —
✅ 111_pod_names_contain_service 26.8s 4 8 $0.2128 80,987 79,249 22,115 1,738 826 56,556 22,693 — —
✅ 112_find_pvcs_by_uuid 20.5s 4 4 $0.2006 81,822 80,589 22,507 1,233 581 58,070 22,519 — —
✅ 12_job_crashing 43.0s 7 18 $0.3304 176,769 173,846 28,874 2,923 687 144,003 29,843 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 49.2s 7 15 $0.3150 161,701 158,730 27,640 2,971 955 130,560 28,170 — —
✅ 227_count_configmaps_per_namespace[0] 25.0s 6 11 $0.2192 117,173 115,635 21,204 1,538 519 94,214 21,421 — —
✅ 24_misconfigured_pvc 36.1s 6 14 $0.2642 126,686 124,296 24,096 2,390 725 99,376 24,920 — —
✅ 43_current_datetime_from_prompt 4.2s 1 — $0.1114 17,455 17,334 17,334 121 121 0 17,334 — —
✅ 61_exact_match_counting 9.6s 3 2 $0.1373 53,323 53,027 17,978 296 158 35,038 17,989 — —
Total 29.9s avg 5.1 avg 10.7 avg $2.5086 1,113,229 1,094,735 29,462 18,494 955 832,884 261,851 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 Run @ 9f74366 (#23141240056)

✅ Results of HolmesGPT evals

Automatically triggered by commit 9f74366 on branch claude/tdd-mcp-json-schemas-4DkIW

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 26.9s 4 8 $0.2181 82,161 80,357 22,649 1,804 892 57,127 23,230 — —
✅ 101_loki_historical_logs_pod_deleted 59.2s 9 17 $0.3603 210,868 207,294 28,090 3,574 520 177,921 29,373 — —
✅ 111_pod_names_contain_service 31.9s 5 9 $0.2344 103,167 101,154 22,766 2,013 569 77,815 23,339 — —
✅ 112_find_pvcs_by_uuid 22.8s 5 4 $0.2094 99,928 98,714 22,546 1,214 337 76,155 22,559 — —
✅ 12_job_crashing 29.3s 5 11 $0.2438 109,401 107,540 24,202 1,861 496 82,374 25,166 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 42.6s 6 15 $0.3024 136,578 133,541 27,486 3,037 986 105,695 27,846 — —
✅ 227_count_configmaps_per_namespace[0] 25.6s 6 10 $0.2187 115,950 114,425 21,028 1,525 515 92,861 21,564 — —
✅ 24_misconfigured_pvc 35.9s 6 15 $0.2691 127,835 125,418 24,414 2,417 899 99,763 25,655 — —
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1120 17,482 17,334 17,334 148 148 0 17,334 — —
✅ 61_exact_match_counting 11.2s 3 2 $0.1386 53,449 53,118 18,037 331 184 35,070 18,048 — —
Total 29.0s avg 5.0 avg 10.1 avg $2.3067 1,056,819 1,038,895 28,090 17,924 986 804,781 234,114 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 45b8578 on branch claude/tdd-mcp-json-schemas-4DkIW (labels: evals-tag-mcp)

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 10/10 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.5s 4 10 $0.2305 83,252 81,146 23,277 2,106 980 57,104 24,042 — —
✅ 101_loki_historical_logs_pod_deleted 48.9s 6 13 $0.3035 138,990 136,016 26,933 2,974 866 107,694 28,322 — —
✅ 111_pod_names_contain_service 42.2s 7 11 $0.2644 144,073 141,748 23,540 2,325 608 118,193 23,555 — —
✅ 112_find_pvcs_by_uuid 17.2s 3 3 $0.1742 59,962 59,085 21,335 877 457 37,739 21,346 — —
✅ 12_job_crashing 33.9s 5 12 $0.2554 111,880 109,880 24,896 2,000 604 83,363 26,517 — —
✅ 176_network_policy_blocking_traffic_no_runbooks 50.0s 7 14 $0.2962 156,281 153,669 26,188 2,612 607 126,708 26,961 — —
✅ 227_count_configmaps_per_namespace[0] 24.5s 5 9 $0.2017 95,065 93,802 20,926 1,263 591 72,242 21,560 — —
✅ 24_misconfigured_pvc 31.8s 5 11 $0.2280 100,833 98,918 22,008 1,915 615 75,988 22,930 — —
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1091 17,072 16,945 16,945 127 127 0 16,945 — —
✅ 61_exact_match_counting 11.3s 3 3 $0.1388 52,903 52,530 17,936 373 226 34,583 17,947 — —
Total 29.6s avg 4.6 avg 9.6 avg $2.2017 960,311 943,739 26,933 16,572 980 713,614 230,125 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tdd-mcp-json-schemas-4DkIW -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/tdd-mcp-json-schemas-4DkIW -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for faa1b834 (built in 1m 31s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:faa1b834
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:faa1b834 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:faa1b834
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:faa1b834
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:faa1b834
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:faa1b834 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:faa1b834
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:faa1b834

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:faa1b834 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:faa1b834

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:faa1b834 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:faa1b834

@coderabbitai

coderabbitai Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Replaces model-based strict tool-call control with a global STRICT_TOOL_CALLS_ENABLED (inverted HOLMES_DISABLE_STRICT_TOOL_CALLS), adds TOOL_SCHEMA_NO_PARAM_OBJECT_IF_NO_PARAMS, introduces per-tool strict-compatibility checks, preserves dynamic-key additionalProperties and JSON Schema validation keywords, adds parameter coercion, and removes model arg from OpenAI formatting APIs.

Changes

Cohort / File(s) Summary
Docs
docs/reference/environment-variables.md
Renamed LLMS_WITH_STRICT_TOOL_CALLS → HOLMES_DISABLE_STRICT_TOOL_CALLS, updated default/semantics, added TOOL_SCHEMA_NO_PARAM_OBJECT_IF_NO_PARAMS and examples.
Env config
holmes/common/env_vars.py
Removed LLMS_WITH_STRICT_TOOL_CALLS; added STRICT_TOOL_CALLS_ENABLED computed as the negation of HOLMES_DISABLE_STRICT_TOOL_CALLS.
Schema formatting
holmes/core/openai_formatting.py
Replaced model-based strictness with STRICT_TOOL_CALLS_ENABLED + per-tool compatibility check (_is_tool_strict_compatible); refined object/array handling of additionalProperties/items and merged passthrough JSON Schema keywords; removed target_model from formatter API.
Tool models & invocation
holmes/core/tools.py
Added additional_properties and json_schema_extra to ToolParameter; added is_strict_compatible() and primary_type; added _coerce_params() with coercion via coerce_params; changed get_openai_format() signature and coercion of params before invoke.
Toolset parsing (MCP)
holmes/plugins/toolsets/mcp/toolset_mcp.py
Parse and resolve additionalProperties (bool or schema, including $ref), preserve anyOf/oneOf, capture validation keywords into json_schema_extra, and pass through to ToolParameter.
Tool executor & callers
holmes/core/tools_utils/tool_executor.py, holmes/core/tool_calling_llm.py
Removed target_model param from get_all_tools_openai_format and its call sites; tools produce OpenAI format independent of target model.
Coercion utility
holmes/core/json_schema_coerce.py
New module: coerce_params(...) to coerce LLM-provided params to schema types with optional strict mode and behavior for stringified JSON/array wrapping and conservative scalar coercions.
Tests
tests/test_openai_formatting.py, tests/core/test_todo_write_tool.py, tests/test_mcp_toolset.py, tests/test_json_schema_coerce.py
Added/updated tests for strict-mode behavior, strict-compatibility, additionalProperties/json_schema_extra passthrough, parameter coercion, and updated usages to get_openai_format() without model arg.

Sequence Diagram(s)

sequenceDiagram
    participant Env as Environment
    participant EnvVars as holmes/common/env_vars.py
    participant Executor as tool_executor.py
    participant Tools as holmes/core/tools.py
    participant Formatting as holmes/core/openai_formatting.py

    Env->>EnvVars: set HOLMES_DISABLE_STRICT_TOOL_CALLS
    EnvVars->>EnvVars: STRICT_TOOL_CALLS_ENABLED = NOT(HOLMES_DISABLE_STRICT_TOOL_CALLS)

    Executor->>Formatting: request all tool schemas
    Formatting->>EnvVars: read STRICT_TOOL_CALLS_ENABLED
    Formatting->>Tools: call tool.get_openai_format()
    Tools->>Tools: is_strict_compatible(parameters)
    Tools-->>Formatting: return parameters + compatibility info

    Formatting->>Formatting: strict_mode = STRICT_TOOL_CALLS_ENABLED && compatible
    alt strict_mode true
        Formatting->>Formatting: enforce required + additionalProperties: false
    else strict_mode false
        Formatting->>Formatting: preserve dynamic-key additionalProperties and json_schema_extra
    end

    alt TOOL_SCHEMA_NO_PARAM_OBJECT_IF_NO_PARAMS true
        Formatting->>Formatting: omit empty "parameters" object for paramless tools (rgba(0,128,0,0.5))
    end

    Formatting-->>Executor: return formatted OpenAI tool schemas
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • arikalon1
🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 30.59% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately and specifically describes the main change: enabling strict tool calling universally (by default) with per-tool compatibility checks.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 15, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 45b8578
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69b83a4380a11000089390de
😎 Deploy Preview https://deploy-preview-1790--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Mar 15, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.45s 12.04s -4.9%
Warm Mean 5.40s 5.27s +2.4%
Warm Min 5.32s 5.26s
Warm Max 5.47s 5.28s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 19.81s 28.54s -30.6%
Warm Mean 8.93s 9.11s -2.0%
Warm Min 8.75s 8.40s
Warm Max 9.10s 9.89s

PR: faa1b834 | Master: 3d97a980 | Iterations: 5

Now that strict mode is universal (not model-dependent), the target_model
parameter is no longer used in format_tool_to_open_ai_standard or any of
its callers.

https://claude.ai/code/session_01UJPFbZ7Y33QNGRHFDP4sRr
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/core/openai_formatting.py (1)

37-54: ⚠️ Potential issue | 🟡 Minor

Bug: additionalProperties schema lost when object has both properties and additionalProperties.

When an object has explicit properties AND an additionalProperties schema (e.g., {"properties": {"name": {...}}, "additionalProperties": {"type": "string"}}), the current elif structure on Line 51 is never reached because Line 41 evaluates to True first.

With strict_mode=False (due to dynamic keys), the schema becomes {"type": "object", "properties": {...}} without the additionalProperties, losing the dynamic-key schema.

🐛 Proposed fix to preserve additionalProperties alongside properties
     if param_type == "object":
         type_obj = {"type": "object"}
 
         # Use explicit properties if provided
         if hasattr(param_attributes, "properties") and param_attributes.properties:
             type_obj["properties"] = {
                 name: type_to_open_ai_schema(prop, strict_mode)
                 for name, prop in param_attributes.properties.items()
             }
             if strict_mode:
                 type_obj["required"] = list(param_attributes.properties.keys())
                 type_obj["additionalProperties"] = False
+            elif hasattr(param_attributes, "additional_properties") and param_attributes.additional_properties not in (None, False):
+                type_obj["additionalProperties"] = param_attributes.additional_properties
 
         # Preserve additionalProperties schema for dynamic-key objects
         elif hasattr(param_attributes, "additional_properties") and param_attributes.additional_properties not in (None, False):
             type_obj["additionalProperties"] = param_attributes.additional_properties
         elif strict_mode:
             type_obj["additionalProperties"] = False
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/openai_formatting.py` around lines 37 - 54, The object schema
builder (in the block handling param_type == "object", involving
param_attributes, type_obj, type_to_open_ai_schema and strict_mode) currently
ignores param_attributes.additional_properties when properties exist because of
the if/elif chain; change the logic so that after populating
type_obj["properties"] you also check for a non-None/non-False
param_attributes.additional_properties and set type_obj["additionalProperties"]
to that value (but if strict_mode is True, keep setting
type_obj["additionalProperties"] = False), ensuring additionalProperties is
preserved alongside properties rather than skipped.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@holmes/core/openai_formatting.py`:
- Around line 37-54: The object schema builder (in the block handling param_type
== "object", involving param_attributes, type_obj, type_to_open_ai_schema and
strict_mode) currently ignores param_attributes.additional_properties when
properties exist because of the if/elif chain; change the logic so that after
populating type_obj["properties"] you also check for a non-None/non-False
param_attributes.additional_properties and set type_obj["additionalProperties"]
to that value (but if strict_mode is True, keep setting
type_obj["additionalProperties"] = False), ensuring additionalProperties is
preserved alongside properties rather than skipped.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 5d8818a9-c260-4f48-8bca-daefc22bb8e9

📥 Commits

Reviewing files that changed from the base of the PR and between 9387f4c and bf606bf.

📒 Files selected for processing (6)
  • docs/reference/environment-variables.md
  • holmes/common/env_vars.py
  • holmes/core/openai_formatting.py
  • holmes/core/tools.py
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
  • tests/test_openai_formatting.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
holmes/core/openai_formatting.py (2)

115-121: ⚠️ Potential issue | 🟠 Major

Add None to enum values whenever the schema declares nullability, not just in strict optional mode.

When a parameter has an enum and the source schema is nullable (e.g., ["string", "null"]), line 86 wraps the type with anyOf to allow null, but lines 115-121 only add None to enum values for strict optional parameters. This creates a schema that advertises nullability via the type but rejects null via enum constraints.

Suggested fix
         if hasattr(param_attributes, "enum") and param_attributes.enum:
             enum_values = list(
                 param_attributes.enum
             )  # Create a copy to avoid modifying original
-            # In strict mode, optional parameters need None in their enum to match the type allowing null
+            # Add None to enum whenever schema allows null
+            is_nullable_from_schema = (
+                isinstance(param_attributes.type, list)
+                and any(t == "null" for t in param_attributes.type)
+            )
             if (
-                strict_mode
-                and not param_attributes.required
+                (is_nullable_from_schema or (strict_mode and not param_attributes.required))
                 and None not in enum_values
             ):
                 enum_values.append(None)
             tool_properties[param_name]["enum"] = enum_values
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/openai_formatting.py` around lines 115 - 121, The enum list is
only appended with None when strict_mode and the parameter is optional, causing
a mismatch when the schema itself is nullable; update the logic around
enum_values and tool_properties[param_name]["enum"] so that if the source schema
declares nullability (e.g., param_attributes indicates nullable or its type
includes "null") you append None to enum_values regardless of strict_mode or
param_attributes.required, ensuring enum constraints match the nullable anyOf
wrapper.

41-55: ⚠️ Potential issue | 🟠 Major

Decouple additionalProperties handling from the properties branch.

The current if/elif structure at lines 41-55 causes additional_properties to be silently dropped whenever explicit properties are defined. In non-strict mode, an object with both fixed keys (properties) and a schema for dynamic keys (additional_properties) loses the dynamic-key rule. This breaks mixed schemas per JSON Schema semantics, which explicitly support both simultaneously.

The fix decouples the logic: handle properties independently, then check additional_properties as a separate concern, allowing both to coexist.

Suggested fix
     if param_type == "object":
         type_obj = {"type": "object"}

         # Use explicit properties if provided
         if hasattr(param_attributes, "properties") and param_attributes.properties:
             type_obj["properties"] = {
                 name: type_to_open_ai_schema(prop, strict_mode)
                 for name, prop in param_attributes.properties.items()
             }
             if strict_mode:
                 type_obj["required"] = list(param_attributes.properties.keys())
-                type_obj["additionalProperties"] = False
-
-        # Preserve additionalProperties schema for dynamic-key objects
-        elif hasattr(param_attributes, "additional_properties") and param_attributes.additional_properties not in (None, False):
-            type_obj["additionalProperties"] = param_attributes.additional_properties
-        elif strict_mode:
-            type_obj["additionalProperties"] = False
+
+        additional_properties = (
+            param_attributes.additional_properties
+            if hasattr(param_attributes, "additional_properties")
+            else None
+        )
+        if strict_mode:
+            type_obj["additionalProperties"] = False
+        elif additional_properties is not None:
+            type_obj["additionalProperties"] = additional_properties
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/openai_formatting.py` around lines 41 - 55, The code currently
uses an if/elif that drops param_attributes.additional_properties when
param_attributes.properties exists; change the logic in the block around
type_to_open_ai_schema/param_attributes handling so properties and
additional_properties are handled independently: first, if param_attributes has
properties, set type_obj["properties"] = { ... } and if strict_mode set
type_obj["required"] = list(...) (do not set additionalProperties here), then in
a separate conditional check param_attributes.additional_properties and if it's
not None/False set type_obj["additionalProperties"] =
param_attributes.additional_properties; finally, if no additionalProperties was
set and strict_mode is True set type_obj["additionalProperties"] = False.
Reference symbols: param_attributes, type_obj, type_to_open_ai_schema,
properties, additional_properties, strict_mode.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@holmes/core/openai_formatting.py`:
- Around line 115-121: The enum list is only appended with None when strict_mode
and the parameter is optional, causing a mismatch when the schema itself is
nullable; update the logic around enum_values and
tool_properties[param_name]["enum"] so that if the source schema declares
nullability (e.g., param_attributes indicates nullable or its type includes
"null") you append None to enum_values regardless of strict_mode or
param_attributes.required, ensuring enum constraints match the nullable anyOf
wrapper.
- Around line 41-55: The code currently uses an if/elif that drops
param_attributes.additional_properties when param_attributes.properties exists;
change the logic in the block around type_to_open_ai_schema/param_attributes
handling so properties and additional_properties are handled independently:
first, if param_attributes has properties, set type_obj["properties"] = { ... }
and if strict_mode set type_obj["required"] = list(...) (do not set
additionalProperties here), then in a separate conditional check
param_attributes.additional_properties and if it's not None/False set
type_obj["additionalProperties"] = param_attributes.additional_properties;
finally, if no additionalProperties was set and strict_mode is True set
type_obj["additionalProperties"] = False. Reference symbols: param_attributes,
type_obj, type_to_open_ai_schema, properties, additional_properties,
strict_mode.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 3298a8b6-c6ae-4474-9488-72cf0930d235

📥 Commits

Reviewing files that changed from the base of the PR and between bf606bf and 4876786.

📒 Files selected for processing (6)
  • holmes/core/openai_formatting.py
  • holmes/core/tool_calling_llm.py
  • holmes/core/tools.py
  • holmes/core/tools_utils/tool_executor.py
  • tests/core/test_todo_write_tool.py
  • tests/test_openai_formatting.py
💤 Files with no reviewable changes (1)
  • holmes/core/tool_calling_llm.py
🚧 Files skipped from review as they are similar to previous changes (2)
  • tests/test_openai_formatting.py
  • holmes/core/tools.py

…dation keywords

Two issues fixed:

1. additionalProperties with anyOf/oneOf was being collapsed to just the
   first branch by _resolve_schema (e.g. string|array became just string).
   Now we only resolve $ref inside additionalProperties and preserve
   compound keywords verbatim.

2. JSON Schema validation keywords (minItems, maxItems, minimum, maximum,
   minLength, maxLength, pattern, default) were dropped during MCP schema
   parsing. Added json_schema_extra field to ToolParameter that captures
   these and passes them through to the OpenAI-formatted schema so the
   LLM sees constraints like array length limits.

https://claude.ai/code/session_01UJPFbZ7Y33QNGRHFDP4sRr
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/core/openai_formatting.py (1)

41-55: ⚠️ Potential issue | 🟠 Major

additionalProperties is dropped for object schemas that also define properties.

When properties exists, the current if/elif chain skips explicit additional_properties, so hybrid object schemas lose dynamic-key constraints in generated output.

💡 Proposed fix
     if param_type == "object":
         type_obj = {"type": "object"}

         # Use explicit properties if provided
         if hasattr(param_attributes, "properties") and param_attributes.properties:
             type_obj["properties"] = {
                 name: type_to_open_ai_schema(prop, strict_mode)
                 for name, prop in param_attributes.properties.items()
             }
             if strict_mode:
                 type_obj["required"] = list(param_attributes.properties.keys())
-                type_obj["additionalProperties"] = False
+        # Preserve explicit additionalProperties (including schema dicts) whenever provided
+        if (
+            hasattr(param_attributes, "additional_properties")
+            and param_attributes.additional_properties is not None
+        ):
+            type_obj["additionalProperties"] = param_attributes.additional_properties
+        elif strict_mode:
+            # Strict fallback when source schema didn't specify additionalProperties
+            type_obj["additionalProperties"] = False
-
-        # Preserve additionalProperties schema for dynamic-key objects
-        elif hasattr(param_attributes, "additional_properties") and param_attributes.additional_properties not in (None, False):
-            type_obj["additionalProperties"] = param_attributes.additional_properties
-        elif strict_mode:
-            type_obj["additionalProperties"] = False
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/openai_formatting.py` around lines 41 - 55, The object-schema
branch in type_to_open_ai_schema drops param_attributes.additional_properties
when param_attributes.properties exists; update the logic so
additional_properties is honored whether or not properties are present by
checking param_attributes.additional_properties separately (not in an elif) and
setting type_obj["additionalProperties"] =
param_attributes.additional_properties when it's not None/False, then apply
strict_mode overrides (setting False only if additionalProperties not provided
and strict_mode). Locate the handling around param_attributes, properties, and
additional_properties in openai_formatting.py and refactor the if/elif chain so
both "properties" and "additionalProperties" can be set for the same schema.
🧹 Nitpick comments (1)
tests/test_mcp_toolset.py (1)

582-632: Add a regression test for objects that combine properties and additionalProperties.

This suite currently validates dynamic-key-only objects, but not hybrid schemas (properties + map-style additionalProperties). Adding that case will protect against dropping map constraints during OpenAI schema formatting.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/test_mcp_toolset.py` around lines 582 - 632, Add a regression test
alongside test_additional_properties_anyof_preserved that exercises a hybrid
object schema which defines both explicit properties and additionalProperties
(e.g., a Tool with inputSchema.properties containing a named property like "id"
and additionalProperties with anyOf for dynamic keys); create the RemoteMCPTool
via RemoteMCPTool.create, retrieve the parameters["filters" or the hybrid field]
to assert its additional_properties still equals the anyOf map, then call
get_openai_format() and assert the resulting
openai_format["function"]["parameters"]["properties"][<hybrid_field>] contains
"additionalProperties" with an "anyOf" array of the expected length and elements
so the map-style constraints are not dropped during OpenAI formatting.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@holmes/core/openai_formatting.py`:
- Around line 41-55: The object-schema branch in type_to_open_ai_schema drops
param_attributes.additional_properties when param_attributes.properties exists;
update the logic so additional_properties is honored whether or not properties
are present by checking param_attributes.additional_properties separately (not
in an elif) and setting type_obj["additionalProperties"] =
param_attributes.additional_properties when it's not None/False, then apply
strict_mode overrides (setting False only if additionalProperties not provided
and strict_mode). Locate the handling around param_attributes, properties, and
additional_properties in openai_formatting.py and refactor the if/elif chain so
both "properties" and "additionalProperties" can be set for the same schema.

---

Nitpick comments:
In `@tests/test_mcp_toolset.py`:
- Around line 582-632: Add a regression test alongside
test_additional_properties_anyof_preserved that exercises a hybrid object schema
which defines both explicit properties and additionalProperties (e.g., a Tool
with inputSchema.properties containing a named property like "id" and
additionalProperties with anyOf for dynamic keys); create the RemoteMCPTool via
RemoteMCPTool.create, retrieve the parameters["filters" or the hybrid field] to
assert its additional_properties still equals the anyOf map, then call
get_openai_format() and assert the resulting
openai_format["function"]["parameters"]["properties"][<hybrid_field>] contains
"additionalProperties" with an "anyOf" array of the expected length and elements
so the map-style constraints are not dropped during OpenAI formatting.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: bdc7a32f-78c5-4bbc-afdd-8916e9d5173e

📥 Commits

Reviewing files that changed from the base of the PR and between 4876786 and 9f74366.

📒 Files selected for processing (4)
  • holmes/core/openai_formatting.py
  • holmes/core/tools.py
  • holmes/plugins/toolsets/mcp/toolset_mcp.py
  • tests/test_mcp_toolset.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • holmes/plugins/toolsets/mcp/toolset_mcp.py

claude added 2 commits March 16, 2026 12:16
LLMs sometimes send serialized JSON strings for array/object parameters
(e.g. '["cpu"]' instead of ["cpu"]), especially when strict mode is
disabled due to dynamic-key params. This adds a _coerce_params() method
on Tool that detects type mismatches against the schema and parses
stringified values before forwarding to the underlying tool/MCP server.

Also adds ToolParameter.primary_type property for extracting the
non-null type from nullable type lists.

https://claude.ai/code/session_01UJPFbZ7Y33QNGRHFDP4sRr
Signed-off-by: Claude <noreply@anthropic.com>
… support

Move LLM parameter coercion from Tool._coerce_params into a reusable
holmes.core.json_schema_coerce module. Add coercions for string→int,
string→float, string→bool, and single-value→array wrapping. Include a
strict mode flag that limits coercion to structural fixes only (JSON
parsing + array wrapping) while skipping potentially lossy scalar
conversions. Document why Pydantic TypeAdapter was not used.

https://claude.ai/code/session_01UJPFbZ7Y33QNGRHFDP4sRr
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
tests/test_json_schema_coerce.py (1)

13-14: Add type hints to the helper function.

Per coding guidelines, type hints are required for Python files. The helper function should have proper type annotations.

✨ Suggested improvement
-def _schema(**fields: ToolParameter) -> dict:
-    return fields
+def _schema(**fields: ToolParameter) -> dict[str, ToolParameter]:
+    return fields
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/test_json_schema_coerce.py` around lines 13 - 14, Update the helper
_schema to include proper type annotations: change its signature to accept
keyword args with values of ToolParameter (keep the existing **fields parameter
name) and annotate the return type as a mapping from str to ToolParameter (e.g.,
dict[str, ToolParameter] or typing.Dict[str, ToolParameter] depending on
supported Python version); also add the necessary typing import if not already
present. This targets the function named _schema and the ToolParameter type
referenced in the test.
holmes/core/json_schema_coerce.py (1)

176-232: Consider consistent shallow copy behavior.

The docstring states "returns a shallow copy" but line 208 returns the original params when it's empty or schema is empty. While this has no practical impact (empty dicts have nothing to mutate), it's a minor documentation inconsistency.

✨ Suggested fix for consistency
     if not schema or not params:
-        return params
+        return dict(params) if params else {}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/json_schema_coerce.py` around lines 176 - 232, The early-return
in coerce_params currently returns the original params when params or schema are
falsy, which contradicts the docstring promise of returning a shallow copy;
change the branch guarded by "if not schema or not params" to return a shallow
copy (e.g., dict(params)) instead of returning params directly so coerce_params
always returns a new dict object while preserving behavior for empty inputs.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@holmes/core/json_schema_coerce.py`:
- Around line 176-232: The early-return in coerce_params currently returns the
original params when params or schema are falsy, which contradicts the docstring
promise of returning a shallow copy; change the branch guarded by "if not schema
or not params" to return a shallow copy (e.g., dict(params)) instead of
returning params directly so coerce_params always returns a new dict object
while preserving behavior for empty inputs.

In `@tests/test_json_schema_coerce.py`:
- Around line 13-14: Update the helper _schema to include proper type
annotations: change its signature to accept keyword args with values of
ToolParameter (keep the existing **fields parameter name) and annotate the
return type as a mapping from str to ToolParameter (e.g., dict[str,
ToolParameter] or typing.Dict[str, ToolParameter] depending on supported Python
version); also add the necessary typing import if not already present. This
targets the function named _schema and the ToolParameter type referenced in the
test.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: ccd7cd52-a774-4e3c-b561-e9bfdaf6aead

📥 Commits

Reviewing files that changed from the base of the PR and between 38cbc93 and 68b14da.

📒 Files selected for processing (4)
  • holmes/core/json_schema_coerce.py
  • holmes/core/tools.py
  • tests/test_json_schema_coerce.py
  • tests/test_openai_formatting.py

@aantn
aantn enabled auto-merge (squash) March 16, 2026 17:13
@aantn
aantn merged commit 0b4d043 into master Mar 16, 2026
21 of 22 checks passed
@aantn
aantn deleted the claude/tdd-mcp-json-schemas-4DkIW branch March 16, 2026 17:17
henrikrexed pushed a commit to henrikrexed/holmesgpt that referenced this pull request Mar 17, 2026
…ecks (HolmesGPT#1790)

## Summary
This PR refactors strict tool calling from a model-based allowlist to a
universal default with automatic per-tool compatibility detection.
Strict mode is now enabled by default for all models and can only be
disabled globally via environment variable. Tools with dynamic-key
parameters are automatically excluded from strict mode on a per-tool
basis.

## Key Changes

- **Simplified strict mode configuration**: Replaced
`LLMS_WITH_STRICT_TOOL_CALLS` model allowlist with
`HOLMES_DISABLE_STRICT_TOOL_CALLS` boolean flag. Strict mode is now
enabled universally by default.

- **Added `is_strict_compatible()` method to `ToolParameter`**:
Recursively checks if a parameter and all nested parameters can be used
in strict mode. Parameters with dynamic keys (`additionalProperties` set
to a schema or `True`) are marked as incompatible.

- **Implemented `_is_tool_strict_compatible()` helper**: Validates all
parameters in a tool before enabling strict mode. Tools with any
incompatible parameters are automatically excluded from strict mode.

- **Improved `additionalProperties` handling in schema generation**:
- Objects with explicit properties now set `additionalProperties: false`
only in strict mode
- Objects with dynamic-key schemas preserve the `additionalProperties`
schema definition
- Prevents strict mode from being applied to tools that require dynamic
keys

- **Added `additional_properties` field to `ToolParameter`**: Stores the
`additionalProperties` JSON Schema value (None, False, or a schema dict)
for proper schema preservation and strict mode compatibility checking.

- **Updated MCP toolset parser**: Now extracts and preserves
`additionalProperties` from MCP tool schemas, resolving any nested
`$ref` or `anyOf` references.

- **Updated documentation**: Clarified that strict mode is now universal
with automatic per-tool exclusions for dynamic-key parameters.

## Implementation Details

The strict mode logic now follows this flow:
1. Check if `STRICT_TOOL_CALLS_ENABLED` is true (default unless
`HOLMES_DISABLE_STRICT_TOOL_CALLS=true`)
2. For each tool, verify all parameters are strict-compatible via
`_is_tool_strict_compatible()`
3. Only enable strict mode if both conditions are met
4. When generating schemas in strict mode, set `additionalProperties:
false` only on objects with explicit properties, preserving dynamic-key
schemas where needed

This ensures OpenAI and Anthropic's strict mode requirements (all
objects must have `additionalProperties: false`) are met while still
supporting tools with dynamic parameters by excluding them from strict
mode.

https://claude.ai/code/session_01UJPFbZ7Y33QNGRHFDP4sRr

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* New env var TOOL_SCHEMA_NO_PARAM_OBJECT_IF_NO_PARAMS (default false)
to omit empty parameter objects.
* Automatic coercion of tool-call parameters to match declared JSON
Schema types (e.g., parsing stringified JSON, wrapping single values
into arrays).

* **Bug Fixes**
* Strict tool-calling is enabled by default; disable with
HOLMES_DISABLE_STRICT_TOOL_CALLS.
* Strict-mode now respects per-tool exceptions for dynamic-key or nested
parameter schemas.
  * Tool OpenAI-format generation no longer depends on a target model.

* **Documentation**
* Env var docs updated; LLMS_WITH_STRICT_TOOL_CALLS retired and replaced
by HOLMES_DISABLE_STRICT_TOOL_CALLS.

* **Tests**
  * Added extensive tests for coercion and strict-mode behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants