Skip to content

test(mcp): add property tests for compare_results scoring (task 6.2) - #239

Closed
tonythethompson wants to merge 1 commit into
feat/agent-compare-tool-v2from
test/compare-props-v2
Closed

tonythethompson wants to merge 1 commit into
feat/agent-compare-tool-v2from
test/compare-props-v2

Conversation

@tonythethompson

@tonythethompson tonythethompson commented Aug 10, 2026 •

Copy link
Copy Markdown
Owner

Property 6: Preference-Weighted Scoring (validates Req 7.3, 7.8)

  • 5 hypothesis property tests verifying _normalize_and_score:
    • winner always has highest score
    • all scores in [0.0, 1.0]
    • balanced preference gives equal weights
    • specific preferences still produce valid scores
    • scoring preserves job IDs
  • 3 integration property tests via full compare_results (mocked studio):
    • winner has highest score
    • scores in valid range
    • JSON round-trip
  • 4 direct weight verification example tests:
    • latency preference doubles latency weight
    • size preference doubles size weight
    • accuracy preference doubles accuracy weight
    • balanced weights symmetric

Uses @settings(max_examples=100) per design spec.

Review in cubic

Property 6: Preference-Weighted Scoring (validates Req 7.3, 7.8)

- 5 hypothesis property tests verifying _normalize_and_score:
  - winner always has highest score
  - all scores in [0.0, 1.0]
  - balanced preference gives equal weights
  - specific preferences still produce valid scores
  - scoring preserves job IDs
- 3 integration property tests via full compare_results (mocked studio):
  - winner has highest score
  - scores in valid range
  - JSON round-trip
- 4 direct weight verification example tests:
  - latency preference doubles latency weight
  - size preference doubles size weight
  - accuracy preference doubles accuracy weight
  - balanced weights symmetric

Uses @settings(max_examples=100) per design spec.
@vercel

vercel Bot commented Aug 10, 2026

Copy link
Copy Markdown

Deployment failed for project olive-studio with the following error:

Resource is limited - try again in 1 day (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/trackdub?upgradeToPro=build-rate-limit

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @tonythethompson, you have reached your weekly rate limit of 500000 diff characters.

Please try again later or upgrade to continue using Sourcery

@qodo-code-review

Copy link
Copy Markdown
Contributor

PR Summary by Qodo

test(mcp): property tests for compare_results preference-weighted scoring

🧪 Tests 🕐 40+ Minutes

Grey Divider

AI Description

• Add Hypothesis property tests for preference-weighted scoring in _normalize_and_score.
• Add integration property tests for compare_results with mocked Studio responses.
• Add example tests verifying 2x weights for latency/size/accuracy and balanced.
Diagram

sequenceDiagram
  actor T as "Test suite"
  participant N as "_normalize_and_score"
  participant C as "compare_results"
  participant R as "studio_request (mocked)"
  participant S as "Olive Studio API"

  T->>N: Score jobs with preference
  N-->>T: Scored list (0..1)

  T->>C: compare_results(job_ids, preference)
  C->>R: GET /status/{job_id}
  R-->>C: Completed job + metrics
  C->>N: Normalize + weight + score
  N-->>C: Scored comparison + winner
  C-->>T: Result JSON (round-trip safe)
Loading
High-Level Assessment

The approach (property tests on the pure scoring function plus integration properties on compare_results with a mocked Studio boundary) is the most robust way to validate preference-weighted scoring without coupling tests to external services. Alternatives like only example-based tests or only integration tests would either reduce coverage of edge cases or increase flakiness and runtime.

Files changed (1) +344 / -0

Tests (1) +344 / -0
test_agent_compare_props.pyAdd property/integration tests for preference-weighted scoring +344/-0

Add property/integration tests for preference-weighted scoring

• Introduces Hypothesis strategies and property tests to validate '_normalize_and_score' invariants (winner ordering, score bounds, job ID preservation) across preferences. Adds integration property tests for 'compare_results' using a patched 'studio_request', including JSON round-trip validation. Includes deterministic example tests that assert the intended 2x weighting behavior for latency/size/accuracy and symmetry under balanced scoring.

olive-mcp-server/tests/test_agent_compare_props.py

@qodo-code-review

Copy link
Copy Markdown
Contributor

Code Review by Qodo

🐞 Bugs (2) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Remediation recommended

1. Preference properties do not verify weights 🐞 Bug ⚙ Maintainability
Description
test_balanced_gives_equal_weights and test_specific_preference_increases_metric_contribution
only check score type/range, never comparing results with an independently calculated weighted score
or preference-specific ordering. An implementation that ignored preference could pass these
property tests, leaving the claimed property coverage incomplete despite the separate deterministic
examples.
Code

olive-mcp-server/tests/test_agent_compare_props.py[R128-132]

+        # Re-compute manually with all weights = 1.0
+        # Since balanced already uses 1.0 for all, scores should match.
+        for entry in scored_balanced:
+            assert isinstance(entry["score"], float)
+            assert 0.0 <= entry["score"] <= 1.0
Relevance

●●● Strong

Repo reviewers often accept strengthening tests to assert real logic/determinism, not just
ranges/types.

PR-#75
PR-#198

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The property assertions at these lines only establish type and range. The implementation's
preference-dependent weights are selected at agent_compare.py lines 96-108, but the tests never
assert those weights affect the result; the deterministic tests at lines 253-328 provide only
partial fixed-case coverage.

olive-mcp-server/tests/test_agent_compare_props.py[115-155]
olive-mcp-server/tests/test_agent_compare_props.py[253-328]
olive-mcp-server/olive_mcp_server/tools/agent_compare.py[96-146]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The balanced and specific-preference property tests do not verify weighting; they only assert that output scores are in range.

## Issue Context
The direct examples provide partial coverage, but the generated properties should validate the scoring formula across arbitrary metric sets. Compute min/max normalization independently in the test, apply the expected 1x/2x weights while skipping missing metrics, and compare expected scores with the function output using an appropriate floating-point tolerance.

## Fix Focus Areas
- olive-mcp-server/tests/test_agent_compare_props.py[115-155]
- olive-mcp-server/olive_mcp_server/tools/agent_compare.py[96-146]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Winner property is tautological 🐞 Bug ⚙ Maintainability
Description
The direct property test computes both max_score and winner from the same returned list, so its
assertion passes for every nonempty output and does not validate scoring or winner selection. The
integration test separately checks that compare_results reports an entry with the highest returned
score, but neither test independently derives the expected winner from the input metrics.
Code

olive-mcp-server/tests/test_agent_compare_props.py[R93-97]

+        # Find the max score
+        max_score = max(entry["score"] for entry in scored)
+        winner = max(scored, key=lambda x: x["score"])
+
+        assert winner["score"] == max_score
Relevance

●●● Strong

Team has accepted rewriting tests that don’t actually validate behavior (tautological/doesn’t
exercise real path).

PR-#73

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Lines 94-95 derive both sides of the assertion from scored, so the equality cannot fail for a
nonempty list. The integration assertion at lines 200-207 is meaningful for consistency of the
returned winner, but it still uses returned scores rather than an independently expected winner.

olive-mcp-server/tests/test_agent_compare_props.py[80-97]
olive-mcp-server/tests/test_agent_compare_props.py[176-207]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The direct winner property selects the maximum result and then asserts that it is the maximum, making the property vacuous.

## Issue Context
Retain the integration invariant if useful, but independently calculate expected normalized weighted scores from `metrics_list`, identify the expected job ID, and compare it with the scored output. Define the expected tie behavior if ties are possible.

## Fix Focus Areas
- olive-mcp-server/tests/test_agent_compare_props.py[80-97]
- olive-mcp-server/tests/test_agent_compare_props.py[176-207]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context used
✅ Compliance rules (platform): 25 rules
✅ REVIEW.md
Review mode: 🚀 Fast: This is a self-contained test-only addition for one scoring path, with no runtime, security, API, schema, or other high-blast-radius changes.

Grey Divider

Tip of the day
💡 Did you know, you can reply 'qodo' on any finding to push back, ask questions, or dig deeper

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment on lines +128 to +132
# Re-compute manually with all weights = 1.0
# Since balanced already uses 1.0 for all, scores should match.
for entry in scored_balanced:
assert isinstance(entry["score"], float)
assert 0.0 <= entry["score"] <= 1.0

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

1. Preference properties do not verify weights 🐞 Bug ⚙ Maintainability

test_balanced_gives_equal_weights and test_specific_preference_increases_metric_contribution
only check score type/range, never comparing results with an independently calculated weighted score
or preference-specific ordering. An implementation that ignored preference could pass these
property tests, leaving the claimed property coverage incomplete despite the separate deterministic
examples.
Agent Prompt
## Issue description
The balanced and specific-preference property tests do not verify weighting; they only assert that output scores are in range.

## Issue Context
The direct examples provide partial coverage, but the generated properties should validate the scoring formula across arbitrary metric sets. Compute min/max normalization independently in the test, apply the expected 1x/2x weights while skipping missing metrics, and compare expected scores with the function output using an appropriate floating-point tolerance.

## Fix Focus Areas
- olive-mcp-server/tests/test_agent_compare_props.py[115-155]
- olive-mcp-server/olive_mcp_server/tools/agent_compare.py[96-146]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment on lines +93 to +97
# Find the max score
max_score = max(entry["score"] for entry in scored)
winner = max(scored, key=lambda x: x["score"])

assert winner["score"] == max_score

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

2. Winner property is tautological 🐞 Bug ⚙ Maintainability

The direct property test computes both max_score and winner from the same returned list, so its
assertion passes for every nonempty output and does not validate scoring or winner selection. The
integration test separately checks that compare_results reports an entry with the highest returned
score, but neither test independently derives the expected winner from the input metrics.
Agent Prompt
## Issue description
The direct winner property selects the maximum result and then asserts that it is the maximum, making the property vacuous.

## Issue Context
Retain the integration invariant if useful, but independently calculate expected normalized weighted scores from `metrics_list`, identify the expected job ID, and compare it with the scored output. Define the expected tie behavior if ties are possible.

## Fix Focus Areas
- olive-mcp-server/tests/test_agent_compare_props.py[80-97]
- olive-mcp-server/tests/test_agent_compare_props.py[176-207]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@qodo-code-review

Copy link
Copy Markdown
Contributor

Qodo Fixer

No findings are within the configured fix scope. To change which findings are fixed, adjust the setting on your Qodo configuration page.

@coderabbitai

coderabbitai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@tonythethompson, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 50 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 59eb5259-67c0-4373-b1fb-a7fe9deb4ea2

📥 Commits

Reviewing files that changed from the base of the PR and between 57e1792 and 4299d1a.

📒 Files selected for processing (1)
  • olive-mcp-server/tests/test_agent_compare_props.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@greptile-apps

greptile-apps Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR adds property-based and example tests for preference-weighted comparison scoring, including direct scoring checks and mocked compare_results integration coverage.

  • Exercises winner selection, score bounds, balanced and preference-specific weights, job-ID preservation, and JSON serialization.
  • Configures each property test to run 100 generated examples.

Confidence Score: 4/5

The PR should not merge until Hypothesis is added to the MCP server's development dependencies so CI can collect the new tests.

The new test file imports Hypothesis, while CI installs only the declared dev extra and that extra does not include Hypothesis, causing test collection to fail.

Files Needing Attention: olive-mcp-server/tests/test_agent_compare_props.py; olive-mcp-server/pyproject.toml

Important Files Changed

Filename Overview
olive-mcp-server/tests/test_agent_compare_props.py Adds comprehensive Hypothesis scoring tests, but the unconditional Hypothesis imports cannot be collected in the currently declared CI development environment.

Fix All in Devin

Prompt To Fix All With AI
### Issue 1
olive-mcp-server/tests/test_agent_compare_props.py:18-19
**Undeclared Hypothesis test dependency**

When CI installs the declared `dev` extra and collects this module, the unconditional Hypothesis imports fail with `ModuleNotFoundError`, preventing the MCP test suite from running. Add Hypothesis to the package's development dependencies.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Reviews (1): Last reviewed commit: "test(mcp): add property tests for compar..." | Re-trigger Greptile

Comment on lines +18 to +19
from hypothesis import given, settings
from hypothesis import strategies as st

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Undeclared Hypothesis test dependency

When CI installs the declared dev extra and collects this module, the unconditional Hypothesis imports fail with ModuleNotFoundError, preventing the MCP test suite from running. Add Hypothesis to the package's development dependencies.

Prompt To Fix With AI
This is a comment left during a code review.
Path: olive-mcp-server/tests/test_agent_compare_props.py
Line: 18-19

Comment:
**Undeclared Hypothesis test dependency**

When CI installs the declared `dev` extra and collects this module, the unconditional Hypothesis imports fail with `ModuleNotFoundError`, preventing the MCP test suite from running. Add Hypothesis to the package's development dependencies.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Devin

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4299d1a4f4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +18 to +19
from hypothesis import given, settings
from hypothesis import strategies as st

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Declare Hypothesis in the dev dependencies

The clean python-tests job installs only .[dev] (.github/workflows/ci.yml:116-120), while that extra currently contains only pytest (olive-mcp-server/pyproject.toml:18-21). Consequently this unconditional import raises ModuleNotFoundError: No module named 'hypothesis' during collection and prevents the entire MCP test suite from running; add Hypothesis to the dev extra.

AGENTS.md reference: AGENTS.md:L89-L89

Useful? React with 👍 / 👎.

Comment on lines +152 to +155
# Each preference-weighted scoring must still produce valid scores
for scored in [scored_balanced, scored_latency, scored_size, scored_accuracy]:
for entry in scored:
assert 0.0 <= entry["score"] <= 1.0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Assert the advertised 2x contribution

For generated inputs, this test only checks that the resulting scores remain in range and never compares them with the expected weighted-average formula. For example, changing every preferred weight from 2x to 3x would still pass this property and all three fixed winner examples, despite violating the requirement this test claims to validate; compute the expected normalized score for each preference and assert equality.

Useful? React with 👍 / 👎.

@tonythethompson

Copy link
Copy Markdown
Owner Author

Superseded by consolidated PR #245

@tonythethompson
tonythethompson deleted the test/compare-props-v2 branch August 13, 2026 00:32

This branch was successfully deployed

1 active deployment
Preview — 4299d1a4 Deployed Aug 10, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant