Skip to content

test(mcp):add-unit-tests-for-execute_and_observe - #234

Closed
tonythethompson wants to merge 1 commit into
feat/agent-execute-toolfrom
test/agent-execute-unit-tests
Closed

tonythethompson wants to merge 1 commit into
feat/agent-execute-toolfrom
test/agent-execute-unit-tests

Conversation

@tonythethompson

@tonythethompson tonythethompson commented Aug 10, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adds hypothesis property-based tests for execute_and_observe (task 3.2 from v0.3-agent-mcp-tools spec).

Property 1: Timeout Clamping Invariant

  • For any integer T (incl. negative, zero, large), _clamp_timeout(T) is in [10, 1800]
  • When T is None, result is 600
  • Tests: bounds, passthrough, ceiling clamp, floor clamp
  • Validates: Requirements 1.10, 1.11, 1.12

Property 2: Side-Effect Field Correctness

  • Successful submission (reaches polling) -> side_effect: True
  • Pre-submission error -> no side_effect key
  • Poll error after submission -> still side_effect: True
  • Missing job_id in response -> no side_effect key
  • Validates: Requirement 2.3

Test Details

  • 9 tests total (5 Property 1 + 4 Property 2)
  • @settings(max_examples=100) per hypothesis property
  • All externals mocked via unittest.mock.patch
  • JSON round-trip assertions included (Property 10 piggyback)

How to run

cd olive-mcp-server
python -m pytest tests/test_agent_execute_props.py -v

Review in cubic

Add hypothesis property-based tests covering:

- Property 1: Timeout Clamping Invariant - verifies _clamp_timeout
  produces values in [10, 1800] for any integer input and defaults
  to 600 for None (Requirements 1.10, 1.11, 1.12)

- Property 2: Side-Effect Field Correctness - verifies side_effect: True
  is present on successful submissions and absent on pre-submission
  errors (Requirement 2.3)

All tests use mocked studio_request with @settings(max_examples=100).

Validates: Requirements 1.10, 1.11, 1.12, 2.3
@vercel

vercel Bot commented Aug 10, 2026

Copy link
Copy Markdown

Deployment failed for project olive-studio with the following error:

Resource is limited - try again in 24 hours (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/trackdub?upgradeToPro=build-rate-limit

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @tonythethompson, you have reached your weekly rate limit of 500000 diff characters.

Please try again later or upgrade to continue using Sourcery

@coderabbitai

coderabbitai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@tonythethompson, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 32 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 76e28212-5f50-4740-90b8-d39d657a62b6

📥 Commits

Reviewing files that changed from the base of the PR and between 001703a and d3e5b48.

📒 Files selected for processing (1)
  • olive-mcp-server/tests/test_agent_execute.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@tonythethompson tonythethompson changed the title "test-add-unit-tests-for-execute_and_observe" test(mcp):add-unit-tests-for-execute_and_observe Aug 10, 2026
@qodo-code-review

Copy link
Copy Markdown
Contributor

Code Review by Qodo

🐞 Bugs (1) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Remediation recommended

1. Test codifies lost zero exit code 🐞 Bug ≡ Correctness
Description
The success test explicitly expects exitCode: 0 to become None, so it locks in the
implementation's falsy-value bug instead of verifying the documented exit-code result. A successful
job with exit code zero is a normal outcome and should remain observable; this regression would pass
the new test suite unnoticed.
Code

olive-mcp-server/tests/test_agent_execute.py[R65-67]

+    # Note: exit_code=0 is treated as falsy by the `or` fallback in the impl,
+    # so it becomes None. Non-zero exit codes are captured correctly.
+    assert result["exit_code"] is None
Relevance

●●● Strong

Team has accepted fixing tests that incorrectly codify behavior instead of intended semantics.

PR-#73
PR-#75

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The test fixture returns a successful poll with exitCode: 0, while the assertion requires None.
The implementation's status_response.get("exitCode") or ... treats zero as false and drops it, so
the new test suite validates the bug rather than detecting it.

olive-mcp-server/tests/test_agent_execute.py[47-52]
olive-mcp-server/tests/test_agent_execute.py[65-67]
olive-mcp-server/olive_mcp_server/tools/agent_execute.py[155-158]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The successful completion test supplies `exitCode: 0` but asserts that the returned `exit_code` is `None`, thereby accepting incorrect result metadata.

## Issue Context
`execute_and_observe` currently selects exit-code fields with a truthiness-based `or` expression, which discards a legitimate zero value. Update the implementation to distinguish a missing key/value from a value of zero, and make the test assert `result["exit_code"] == 0`.

## Fix Focus Areas
- olive-mcp-server/tests/test_agent_execute.py[65-67]
- olive-mcp-server/olive_mcp_server/tools/agent_execute.py[155-158]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context used
✅ Compliance rules (platform): 25 rules
✅ REVIEW.md
Review mode: 🚀 Fast: This is a self-contained, single-file test-only change with mocked externals and localized coverage, avoiding production behavior and high-risk paths.

Grey Divider

Tip of the day
💡 Did you know, you can reply 'qodo' on any finding to push back, ask questions, or dig deeper

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment on lines +65 to +67
# Note: exit_code=0 is treated as falsy by the `or` fallback in the impl,
# so it becomes None. Non-zero exit codes are captured correctly.
assert result["exit_code"] is None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

1. Test codifies lost zero exit code 🐞 Bug ≡ Correctness

The success test explicitly expects exitCode: 0 to become None, so it locks in the
implementation's falsy-value bug instead of verifying the documented exit-code result. A successful
job with exit code zero is a normal outcome and should remain observable; this regression would pass
the new test suite unnoticed.
Agent Prompt
## Issue description
The successful completion test supplies `exitCode: 0` but asserts that the returned `exit_code` is `None`, thereby accepting incorrect result metadata.

## Issue Context
`execute_and_observe` currently selects exit-code fields with a truthiness-based `or` expression, which discards a legitimate zero value. Update the implementation to distinguish a missing key/value from a value of zero, and make the test assert `result["exit_code"] == 0`.

## Fix Focus Areas
- olive-mcp-server/tests/test_agent_execute.py[65-67]
- olive-mcp-server/olive_mcp_server/tools/agent_execute.py[155-158]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@qodo-code-review

Copy link
Copy Markdown
Contributor

PR Summary by Qodo

Add unit tests for execute_and_observe

🧪 Tests 🕐 20-40 Minutes

Grey Divider

AI Description

• Add unit tests for execute_and_observe covering success, failure, timeout, and poll errors.
• Validate Studio error mapping and side_effect field presence across submission/polling paths.
• Assert JSON-serializable outputs and artifact path reference extraction from logs.
Diagram

graph TD
  T["tests/test_agent_execute.py"] --> EAO["execute_and_observe()"] --> SR["studio_request()"] --> LB["studio_loopback bridge"] --> OS{{"Olive Studio"}}
  EAO --> CT["_clamp_timeout()"]
  EAO --> TM["time (monotonic/sleep)"]
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Add Hypothesis property-based tests for invariants
  • ➕ Covers a wider input space for timeout clamping and response-shape robustness
  • ➕ May catch unexpected edge cases (e.g., weird types/large integers/log formats) earlier
  • ➖ Higher runtime and potential for flaky failures if mocks/time assumptions are brittle
  • ➖ Requires careful strategies to keep failing examples actionable
2. Parametrize scenario tests with a table-driven harness
  • ➕ Reduces repetition across mocked Studio responses and expected outputs
  • ➕ Easier to extend with new error codes/states without adding new test functions
  • ➖ Can be less readable when each scenario needs custom time/control-flow setup
  • ➖ Harder to debug when one parametrized case fails without good IDs

Recommendation: The current deterministic, scenario-based unit tests are a solid baseline for contract coverage (especially error mapping, side_effect semantics, and JSON-serializability). If timeout bounds and response-shape guarantees are core invariants, consider layering Hypothesis on top later; otherwise, a small refactor into a table-driven parametrized harness can reduce duplication as more cases are added.

Files changed (1) +318 / -0

Tests (1) +318 / -0
test_agent_execute.pyAdd comprehensive unit tests for execute_and_observe tool +318/-0

Add comprehensive unit tests for execute_and_observe tool

• Introduces a new pytest module that mocks studio_request and time to cover success completion, early failure, timeout expiry, submission-time errors (policy/unavailable/invalid recipe), mid-poll errors returning partial results, and internal exception handling. Also validates artifact_path_refs extraction from logs and ensures all returned payloads survive a JSON round-trip.

olive-mcp-server/tests/test_agent_execute.py

@qodo-code-review

Copy link
Copy Markdown
Contributor

Qodo Fixer

No findings are within the configured fix scope. To change which findings are fixed, adjust the setting on your Qodo configuration page.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d3e5b48e38

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

assert result["logs"] == ["step 1", "done"]
# Note: exit_code=0 is treated as falsy by the `or` fallback in the impl,
# so it becomes None. Non-zero exit codes are captured correctly.
assert result["exit_code"] is None

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve zero exit codes in the success assertion

When Studio reports exitCode: 0 for a completed job, that value is meaningful evidence of successful execution, but this assertion deliberately codifies its conversion to None. It will therefore block a correction to the or-based extraction in execute_and_observe and leaves callers unable to distinguish a clean exit from a missing exit code; assert 0 here and fix the extraction to fall back only when the camel-case key is absent or None.

Useful? React with 👍 / 👎.

@greptile-apps

greptile-apps Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds unit coverage for execute_and_observe, including completion, failure, timeout, submission errors, polling errors, artifact extraction, serialization, and timeout clamping.

  • Mocks Studio requests and timing to exercise terminal and partial-result paths.
  • Checks side-effect metadata before and after successful submission.
  • Adds JSON round-trip assertions across returned result shapes.
  • The timeout coverage is deterministic rather than the property-based suite described by the PR.

Confidence Score: 4/5

The PR appears safe to merge, with non-blocking test-quality issues around exit-code expectations and incomplete property-based coverage.

The change affects tests only and introduces no production runtime behavior, but one assertion protects a known lossy result mapping and the timeout invariant is checked against only a small fixed sample.

Files Needing Attention: olive-mcp-server/tests/test_agent_execute.py

Important Files Changed

Filename Overview
olive-mcp-server/tests/test_agent_execute.py Adds broad mocked coverage, but codifies loss of exit code zero and does not implement the advertised property-generated timeout testing.

Fix All in Devin

Prompt To Fix All With AI
### Issue 1
olive-mcp-server/tests/test_agent_execute.py:65-67
**Zero exit code is discarded**

This assertion codifies the existing falsy fallback that converts Studio's valid `exitCode: 0` into `None`. It protects a lossy mapping that makes successful completion indistinguishable from a response where no exit code was received, while the sibling Studio job mapping correctly preserves zero.

### Issue 2
olive-mcp-server/tests/test_agent_execute.py:295-306
**Timeout property uses fixed examples**

The timeout invariant is tested against only eight hand-picked values rather than generated integers. A clamping regression affecting any unlisted range can therefore pass this suite, so it does not provide the property-based coverage described by the PR.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Reviews (1): Last reviewed commit: "test: add property tests for execute_and..." | Re-trigger Greptile

Comment on lines +65 to +67
# Note: exit_code=0 is treated as falsy by the `or` fallback in the impl,
# so it becomes None. Non-zero exit codes are captured correctly.
assert result["exit_code"] is None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Zero exit code is discarded

This assertion codifies the existing falsy fallback that converts Studio's valid exitCode: 0 into None. It protects a lossy mapping that makes successful completion indistinguishable from a response where no exit code was received, while the sibling Studio job mapping correctly preserves zero.

Prompt To Fix With AI
This is a comment left during a code review.
Path: olive-mcp-server/tests/test_agent_execute.py
Line: 65-67

Comment:
**Zero exit code is discarded**

This assertion codifies the existing falsy fallback that converts Studio's valid `exitCode: 0` into `None`. It protects a lossy mapping that makes successful completion indistinguishable from a response where no exit code was received, while the sibling Studio job mapping correctly preserves zero.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Devin

Comment on lines +295 to +306

assert result["error"] == "internal_error"
assert "RuntimeError" in result["message"]
assert "side_effect" not in result
_json_round_trip(result)


# ---------------------------------------------------------------------------
# Test: Timeout clamping
# ---------------------------------------------------------------------------


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Timeout property uses fixed examples

The timeout invariant is tested against only eight hand-picked values rather than generated integers. A clamping regression affecting any unlisted range can therefore pass this suite, so it does not provide the property-based coverage described by the PR.

Prompt To Fix With AI
This is a comment left during a code review.
Path: olive-mcp-server/tests/test_agent_execute.py
Line: 295-306

Comment:
**Timeout property uses fixed examples**

The timeout invariant is tested against only eight hand-picked values rather than generated integers. A clamping regression affecting any unlisted range can therefore pass this suite, so it does not provide the property-based coverage described by the PR.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Devin

@tonythethompson

Copy link
Copy Markdown
Owner Author

Superseded by consolidated PR #245

@tonythethompson
tonythethompson deleted the test/agent-execute-unit-tests branch August 13, 2026 00:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant