Skip to content

feat(mcp): add diagnose_and_fix agent tool - #233

Closed
tonythethompson wants to merge 8 commits into
feat/agent-model-info-toolfrom
test/agent-model-info-unit-tests
Closed

tonythethompson wants to merge 8 commits into
feat/agent-model-info-toolfrom
test/agent-model-info-unit-tests

Conversation

@tonythethompson

@tonythethompson tonythethompson commented Aug 10, 2026 •

Copy link
Copy Markdown
Owner

Implements the diagnose_and_fix MCP tool (agent_diagnosis.py) for the
Phase 3 autonomous agent loop.

  • Validates error_message (1-4000 chars) and recipe (must be dict)
  • Calls troubleshoot_olive_error internally with config context
  • Applies RFC 7386 JSON Merge Patch when KB entry has updated_config
  • Generates human-readable change descriptions
  • Best-effort recipe validation through Studio bridge
  • Maps fix_confidence: high/medium/low/none based on KB match quality
  • Top-level try/except for internal_error safety net
  • Zero new pip dependencies (stdlib only + internal imports)

Requirements: 5.1-5.9, 6.3, 11.1, 11.3, 12.1-12.7

Review in cubic

tonythethompson and others added 7 commits August 10, 2026 02:50
Implements the diagnose_and_fix MCP tool (agent_diagnosis.py) for the
Phase 3 autonomous agent loop.

- Validates error_message (1-4000 chars) and recipe (must be dict)
- Calls troubleshoot_olive_error internally with config context
- Applies RFC 7386 JSON Merge Patch when KB entry has updated_config
- Generates human-readable change descriptions
- Best-effort recipe validation through Studio bridge
- Maps fix_confidence: high/medium/low/none based on KB match quality
- Top-level try/except for internal_error safety net
- Zero new pip dependencies (stdlib only + internal imports)

Requirements: 5.1-5.9, 6.3, 11.1, 11.3, 12.1-12.7
Implements the execute_and_observe MCP tool for the Phase 3 autonomous
agent loop. Submits a recipe to Olive Studio via the loopback bridge,
polls job status at 2-second intervals until a terminal state is reached
or the effective timeout expires.

Key behaviors:
- Timeout clamping: min(max(timeout or 600, 10), 1800)
- Terminal states: completed, failed, cancelled
- Terminal-at-timeout-boundary: terminal wins (timed_out: false)
- Pre-submission errors: no side_effect field
- Post-submission results: side_effect: True
- Logs capped at 200 entries, artifact refs as basenames only
- Top-level try/except for internal_error safety

Requirements: 1.1-1.12, 2.3, 2.4, 2.5, 11.1, 11.3, 12.1-12.7
Implements the compare_results MCP tool (task 6.1) for multi-job
comparison with preference-weighted scoring.

- Validates job_ids count (2-10) and format (^[A-Za-z0-9_-]{1,128}$)
- Normalizes preference (latency/size/accuracy/balanced)
- Fetches job status via studio_request loopback bridge
- Excludes non-terminal, failed, or unfetchable jobs
- Min-max normalizes metrics with lower-is-better inversion
- Applies 2x weight for preferred metric, 1x for others
- Selects highest scored job as winner
- Returns structured comparison with side_effect: False
- Top-level try/except for internal_error safety
- Zero new pip dependencies (stdlib + studio_loopback only)

Requirements: 7.1-7.8, 8.3, 11.1, 11.3, 12.1-12.7
Add hypothesis property-based tests covering:

- Property 1: Timeout Clamping Invariant - verifies _clamp_timeout
  produces values in [10, 1800] for any integer input and defaults
  to 600 for None (Requirements 1.10, 1.11, 1.12)

- Property 2: Side-Effect Field Correctness - verifies side_effect: True
  is present on successful submissions and absent on pre-submission
  errors (Requirement 2.3)

All tests use mocked studio_request with @settings(max_examples=100).

Validates: Requirements 1.10, 1.11, 1.12, 2.3
Add olive-mcp-server/tests/test_agent_model_info.py with 18 pytest unit
tests covering:

- HF API success (params from safetensors and config.num_parameters)
- HF timeout fallback to heuristic (family default, low confidence)
- HF 404 fallback with explicit size token (medium confidence)
- Invalid model_id validation error
- Confidence level differentiation (medium vs low)
- VRAM estimate formula (params_b * 2.0)
- Recommended quantization threshold (int4/int8 at 6B)
- Model type classification via _normalize_model_type
- JSON serialization round-trip for all outputs

All tests monkeypatch _fetch_hf_metadata - no network access.

Requirements: 13.5, 13.6, 13.7
@vercel

vercel Bot commented Aug 10, 2026

Copy link
Copy Markdown

Deployment failed for project olive-studio with the following error:

Resource is limited - try again in 24 hours (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/trackdub?upgradeToPro=build-rate-limit

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @tonythethompson, you have reached your weekly rate limit of 500000 diff characters.

Please try again later or upgrade to continue using Sourcery

@coderabbitai

coderabbitai Bot commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@qodo-code-review[bot], you've reached your PR review limit, so we couldn't start this review.

Next review available in: 41 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 41b058d0-2505-4af7-b440-23502fec72bc

📥 Commits

Reviewing files that changed from the base of the PR and between ddb5de0 and 17a72eb.

📒 Files selected for processing (5)
  • olive-mcp-server/olive_mcp_server/tools/agent_compare.py
  • olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py
  • olive-mcp-server/olive_mcp_server/tools/agent_execute.py
  • olive-mcp-server/tests/test_agent_execute.py
  • olive-mcp-server/tests/test_agent_model_info.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@qodo-code-review

Copy link
Copy Markdown
Contributor

PR Summary by Qodo

Add Phase 3 agent MCP tools (diagnose/execute/compare) with unit tests

✨ Enhancement 🧪 Tests 🕐 40+ Minutes

Grey Divider

AI Description

• Add Phase 3 MCP tools to diagnose errors, execute jobs, and compare completed runs.
• Bridge to Olive Studio for submission/status/validation with consistent error shaping.
• Add unit tests for execute tool and model-info heuristics including JSON round-trip.
Diagram

graph TD
  A["Autonomous agent loop"] --> B["MCP server"] --> C(["execute_and_observe"]) --> F(["studio_request"]) --> G{{"Olive Studio API"}}
  B["MCP server"] --> D(["diagnose_and_fix"]) --> H[("Troubleshooting KB")]
  D(["diagnose_and_fix"]) --> F(["studio_request"]) --> G{{"Olive Studio API"}}
  B["MCP server"] --> E(["compare_results"]) --> F(["studio_request"]) --> G{{"Olive Studio API"}}
  subgraph Legend
    direction LR
    _svc(["Tool / module"]) ~~~ _ext{{"External service"}} ~~~ _db[("Knowledge base")]
  end
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Move comparison/scoring server-side in Studio
  • ➕ Eliminates multiple per-job status calls from MCP tool
  • ➕ Keeps scoring semantics/versioning centralized in Studio
  • ➕ Allows richer metrics and tie-breaking without MCP updates
  • ➖ Requires Studio API changes and deployment coordination
  • ➖ Reduces flexibility to experiment with scoring in MCP layer
2. Use rank-based or z-score normalization for comparison
  • ➕ More robust to outliers than min-max normalization
  • ➕ Better behavior when metrics have tight ranges or skew
  • ➖ Less intuitive to explain/debug than min-max
  • ➖ May be surprising for small job sets (2–3 runs)
3. Validate fixed recipes locally (schema/Pydantic) instead of Studio bridge
  • ➕ Works offline when Studio is unavailable
  • ➕ Deterministic validation without network latency
  • ➖ Adds maintenance burden to keep validation aligned with Studio
  • ➖ May require new dependencies or duplicating Studio rules

Recommendation: The current approach is a reasonable first iteration: keep MCP tools dependency-free and use Studio as the validation authority. If compare_results becomes heavily used or performance-sensitive, consider moving comparison/scoring into a dedicated Studio endpoint. If keeping it in MCP, revisit normalization strategy (rank/z-score) once real-world metric distributions are known.

Files changed (5) +1255 / -0

Enhancement (3) +682 / -0
agent_compare.pyAdd compare_results tool with preference-weighted scoring +279/-0

Add compare_results tool with preference-weighted scoring

• Introduces a Phase 3 tool that fetches multiple job statuses from the Studio bridge, excludes non-completed/unfetchable jobs with reasons, and scores remaining jobs. Implements min-max normalization across latency/size/accuracy and applies a 2x weight to the preferred metric (or balanced). Returns a winner, per-job scores/metrics, and human-readable reasoning with side_effect=False.

olive-mcp-server/olive_mcp_server/tools/agent_compare.py

agent_diagnosis.pyAdd diagnose_and_fix tool with RFC 7386 merge-patch recipe repair +202/-0

Add diagnose_and_fix tool with RFC 7386 merge-patch recipe repair

• Adds a tool that validates inputs, builds a short config_context, and calls troubleshoot_olive_error to find a KB match. When the KB provides an applyable updated_config, applies it via RFC 7386 JSON Merge Patch to produce a fixed recipe and emits human-readable change descriptions. Performs best-effort validation via Studio tool invocation and returns fix_confidence plus side_effect=False with a top-level internal_error safety net.

olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py

agent_execute.pyAdd execute_and_observe tool for job submission and polling +201/-0

Add execute_and_observe tool for job submission and polling

• Implements job submission to Olive Studio and polling status until a terminal state or a clamped timeout is reached. Captures logs (capped), latest metrics, exit code, elapsed time, and extracts artifact basenames from logs; includes side_effect=True for post-submission results and omits side_effect for pre-submission errors. Adds structured handling for common Studio error codes and a poll-error partial-result path.

olive-mcp-server/olive_mcp_server/tools/agent_execute.py

Tests (2) +573 / -0
test_agent_execute.pyAdd unit tests for execute_and_observe behavior and outputs +318/-0

Add unit tests for execute_and_observe behavior and outputs

• Adds coverage for success, early failure, timeout, submission-denied/unavailable/invalid-recipe error mapping, poll errors returning partial results, and internal exception handling. Includes JSON serialization round-trip assertions and verifies artifact basename extraction from logs and timeout clamping.

olive-mcp-server/tests/test_agent_execute.py

test_agent_model_info.pyAdd unit tests for get_model_info HF/heuristic paths +255/-0

Add unit tests for get_model_info HF/heuristic paths

• Adds tests for HuggingFace API success cases (safetensors.total and config.num_parameters) and heuristic fallback behavior on API failure. Verifies confidence levels, derived VRAM/quantization fields, invalid input handling, and JSON round-trip serialization for outputs.

olive-mcp-server/tests/test_agent_model_info.py

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7ca6a12128

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +56 to +59
return {
"latency_ms": _as_float(raw.get("latency_ms")),
"model_size_mb": _as_float(raw.get("model_size_mb")),
"accuracy": _as_float(raw.get("accuracy")),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Do not select a winner from GPU telemetry

For actual Studio jobs, src/server/services/olive/gpu.ts stores {timestamp, gpus} in latestMetrics, and src/server/routes/olive.ts returns that object unchanged; it never contains latency_ms, model_size_mb, or accuracy. Consequently every completed job produces three None values here, all scores become zero, and max() reports the first requested job as the winner despite having no comparison data. Source these measurements from optimization results, or withhold the winner when no requested metric is available.

Useful? React with 👍 / 👎.

Comment on lines +127 to +130
# If Studio is down or returned an error, treat as not validated
if isinstance(response.get("error"), str) and response["error"]:
return False
return True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require a successful validation result

When Studio successfully processes an invalid repaired recipe, validate_optimization_job returns a normal payload with valid: false and an errors list, not an error string. This branch therefore returns True and tells the autonomous retry loop that an invalid fix was validated. Return the payload's valid value after checking for bridge errors.

Useful? React with 👍 / 👎.

Comment on lines +155 to +158
# Capture exit code
exit_code = status_response.get("exitCode") or status_response.get("exit_code")
if exit_code is not None:
last_exit_code = exit_code

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve successful zero exit codes

For every normally completed job, Studio reports exitCode: 0, but the truthiness-based or discards that value and falls through to the absent snake-case field, producing None. The result therefore loses the definitive success code; select the fallback based on key presence or None rather than truthiness.

Useful? React with 👍 / 👎.

@greptile-apps

greptile-apps Bot commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR introduces Phase 3 MCP agent tools for diagnosing and repairing recipes, executing optimization jobs to completion, and comparing completed jobs.

  • Adds knowledge-base-backed diagnosis, JSON Merge Patch repair, confidence mapping, and best-effort Studio validation.
  • Adds job submission, terminal-state polling, partial poll-error results, bounded logs, and artifact-reference extraction.
  • Adds preference-weighted comparison across completed jobs.
  • Adds execution and model-information test coverage.

Confidence Score: 3/5

The PR is not safe to merge until rejected recipe validation is reported accurately and comparison scores use actual optimization-result metrics.

Recipe validation still treats a normal valid=false response as success, while job comparison reads latency, size, and accuracy from a status field that contains only GPU telemetry, collapsing real comparisons to zero-score first-job selection.

Files Needing Attention: olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py; olive-mcp-server/olive_mcp_server/tools/agent_compare.py

Important Files Changed

Filename Overview
olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py Adds troubleshooting-based recipe diagnosis, merge-patch repair, change descriptions, confidence mapping, and Studio validation.
olive-mcp-server/olive_mcp_server/tools/agent_compare.py Adds completed-job filtering, metric normalization, preference weighting, winner selection, and exclusion reasoning.
olive-mcp-server/olive_mcp_server/tools/agent_execute.py Adds Studio job submission, bounded polling, terminal and timeout handling, partial failure results, and artifact-reference extraction.
olive-mcp-server/tests/test_agent_execute.py Covers successful, failed, timed-out, denied, unavailable, malformed, and interrupted execution flows.
olive-mcp-server/tests/test_agent_model_info.py Adds coverage for API-derived and heuristic model metadata, confidence levels, derived recommendations, and serialization.

Reviews (2): Last reviewed commit: "fix: Preserve zero exit codes" | Re-trigger Greptile

Comment on lines +127 to +130
# If Studio is down or returned an error, treat as not validated
if isinstance(response.get("error"), str) and response["error"]:
return False
return True

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Validation result is ignored

When Studio rejects a repaired recipe with valid=false and an errors list, this check sees no top-level error and returns True, causing an invalid recipe to be reported as validated.

Suggested change
# If Studio is down or returned an error, treat as not validated
if isinstance(response.get("error"), str) and response["error"]:
return False
return True
# Require Studio to explicitly confirm that the recipe is valid.
return response.get("valid") is True
Prompt To Fix With AI
This is a comment left during a code review.
Path: olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py
Line: 127-130

Comment:
**Validation result is ignored**

When Studio rejects a repaired recipe with `valid=false` and an `errors` list, this check sees no top-level `error` and returns `True`, causing an invalid recipe to be reported as validated.

```suggestion
    # Require Studio to explicitly confirm that the recipe is valid.
    return response.get("valid") is True
```

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Devin

Comment on lines +57 to +59
"latency_ms": _as_float(raw.get("latency_ms")),
"model_size_mb": _as_float(raw.get("model_size_mb")),
"accuracy": _as_float(raw.get("accuracy")),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Scoring reads GPU telemetry

The Studio status endpoint supplies latestMetrics as GPU telemetry containing timestamp and gpus, but this code reads optimization fields directly from it. Every job therefore receives a zero score, causing the first requested job to be selected regardless of its actual optimization results.

Prompt To Fix With AI
This is a comment left during a code review.
Path: olive-mcp-server/olive_mcp_server/tools/agent_compare.py
Line: 57-59

Comment:
**Scoring reads GPU telemetry**

The Studio status endpoint supplies `latestMetrics` as GPU telemetry containing `timestamp` and `gpus`, but this code reads optimization fields directly from it. Every job therefore receives a zero score, causing the first requested job to be selected regardless of its actual optimization results.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Devin

@qodo-code-review

qodo-code-review Bot commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Code Review by Qodo

🐞 Bugs (1) 📘 Rule violations (1) 📜 Skill insights (0)

Grey Divider


Action required

1. Exit code 0 lost ✓ Resolved 🐞 Bug ≡ Correctness
Description
In execute_and_observe, exitCode: 0 from Studio is treated as falsy due to an or fallback,
causing exit_code to become None for successful jobs. This misreports job outcomes and can break
downstream logic that relies on exit_code being 0 on success.
Code

olive-mcp-server/olive_mcp_server/tools/agent_execute.py[R156-158]

+            exit_code = status_response.get("exitCode") or status_response.get("exit_code")
+            if exit_code is not None:
+                last_exit_code = exit_code
Relevance

●●● Strong

Classic falsy-0 bug; deterministic correctness fix likely accepted.

PR-#8
PR-#156

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Studio’s agent status endpoint returns exitCode in the JSON payload and it can be 0 for success;
the new polling code uses a falsy or fallback that converts 0 into None, and the new unit test
explicitly documents/locks in this incorrect behavior.

olive-mcp-server/olive_mcp_server/tools/agent_execute.py[150-158]
src/server/routes/olive.ts[98-106]
olive-mcp-server/tests/test_agent_execute.py[63-68]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
`execute_and_observe` reads the exit code via `status_response.get("exitCode") or status_response.get("exit_code")`, which drops a valid `0` exit code. This makes successful jobs report `exit_code=None`.

### Issue Context
Studio’s agent status response includes an `exitCode` field (0 for success). The MCP tool should preserve that value exactly.

### Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/agent_execute.py[156-158]

### Suggested fix
Use a `None`-aware fallback instead of `or`, e.g.:
- `exit_code = status_response.get("exitCode")`
- `if exit_code is None: exit_code = status_response.get("exit_code")`
(or `exit_code = status_response.get("exitCode", status_response.get("exit_code"))`).

Update the unit test that currently asserts the buggy behavior so it expects `exit_code == 0`.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Inverted tie scores zero 🐞 Bug ≡ Correctness
Description
In compare_results, when all jobs have the same latency/size value, the code sets normalized=1.0
and then inverts it to 0.0, contradicting the “full score” tie comment. This yields misleading
scores and can skew ranking when some jobs are missing other metrics (e.g., a job with only latency
populated gets 0.0 instead of a tie-best score).
Code

olive-mcp-server/olive_mcp_server/tools/agent_compare.py[R127-130]

+            if mx == mn:
+                normalized = 1.0  # All values equal → full score
+            else:
+                normalized = (val - mn) / (mx - mn)
Relevance

●●● Strong

Tie-handling contradicts comment; team accepts ranking/score correctness fixes.

PR-#75
PR-#73

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The scoring code declares latency/size as inverted metrics, sets normalized=1.0 when mx==mn,
then unconditionally inverts for those metrics, yielding 0.0 for every job on that metric despite
the “full score” tie comment.

olive-mcp-server/olive_mcp_server/tools/agent_compare.py[74-135]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
For inverted metrics (`latency_ms`, `model_size_mb`), the tie branch (`mx == mn`) sets `normalized = 1.0` and then applies inversion, resulting in `0.0` for all jobs on that metric. This contradicts the tie comment and produces misleading scores; it can also affect winner selection when other metrics are missing.

### Issue Context
Latency/size are “lower is better” metrics. If all values are equal, every job should receive an equal *best/tie* contribution for that metric (not zero).

### Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/agent_compare.py[126-134]

### Suggested fix
Handle the `mx == mn` case after considering directionality, e.g. either:
- Skip inversion in the tie case, or
- Set `normalized = 0.0` for inverted metrics when `mx == mn` (so `1 - normalized` becomes `1.0`), and keep `normalized = 1.0` for non-inverted metrics.

Keep behavior consistent with the comment and ensure scores remain meaningful even when only some metrics are present.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Sequential studio_request status fetches 📘 Rule violation ➹ Performance
Description
compare_results fetches each job status sequentially, creating an avoidable request waterfall when
multiple independent job IDs are compared. This can increase overall tool latency proportional to
the number of job IDs.
Code

olive-mcp-server/olive_mcp_server/tools/agent_compare.py[R214-216]

+        for jid in job_ids:
+            response = studio_request("GET", f"{_STATUS_PATH}/{jid}")
+
Relevance

●● Moderate

Parallelizing studio_request is larger change; no close precedent for enforcing parallel fetches
here.

PR-#97
PR-#32

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
PR Compliance ID 2436664 requires avoiding avoidable request waterfalls by running independent
fetches in parallel. In compare_results, each job status is fetched one-by-one in a for loop via
studio_request("GET", ...), even though these requests are independent.

Rule 2436664: Avoid avoidable request waterfalls in data fetching
olive-mcp-server/olive_mcp_server/tools/agent_compare.py[214-216]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`compare_results` currently issues one `studio_request()` per job ID in a loop, which creates an avoidable request waterfall for independent status fetches.

## Issue Context
This tool compares 2–10 jobs; since each status fetch is independent, the total time can be reduced by fetching statuses concurrently (e.g., via a small thread pool for I/O-bound calls, or via a batch endpoint if available).

## Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/agent_compare.py[214-239]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context used
✅ Compliance rules (platform): 25 rules
✅ REVIEW.md
Review mode: ⚖️ Balanced: Downgraded extended -> standard: change is below the extended eligibility bar (hunks 5/18, lines 1255/200; both must reach the floor). Router rationale: This adds substantial runtime agent logic across three independent tools, including side-effectful job execution, diagnosis/recipe patching, and multi-job scoring; the five substantial edit sites create multiple easy-to-miss behavioral defects warranting redundant review.

Grey Divider

Tip of the day
💡 Did you know, you can reply 'qodo' on any finding to push back, ask questions, or dig deeper

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment on lines +214 to +216
for jid in job_ids:
response = studio_request("GET", f"{_STATUS_PATH}/{jid}")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

1. Sequential studio_request status fetches 📘 Rule violation ➹ Performance

compare_results fetches each job status sequentially, creating an avoidable request waterfall when
multiple independent job IDs are compared. This can increase overall tool latency proportional to
the number of job IDs.
Agent Prompt
## Issue description
`compare_results` currently issues one `studio_request()` per job ID in a loop, which creates an avoidable request waterfall for independent status fetches.

## Issue Context
This tool compares 2–10 jobs; since each status fetch is independent, the total time can be reduced by fetching statuses concurrently (e.g., via a small thread pool for I/O-bound calls, or via a batch endpoint if available).

## Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/agent_compare.py[214-239]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

Comment thread olive-mcp-server/olive_mcp_server/tools/agent_execute.py Outdated
Comment on lines +127 to +130
if mx == mn:
normalized = 1.0 # All values equal → full score
else:
normalized = (val - mn) / (mx - mn)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remediation recommended

3. Inverted tie scores zero 🐞 Bug ≡ Correctness

In compare_results, when all jobs have the same latency/size value, the code sets normalized=1.0
and then inverts it to 0.0, contradicting the “full score” tie comment. This yields misleading
scores and can skew ranking when some jobs are missing other metrics (e.g., a job with only latency
populated gets 0.0 instead of a tie-best score).
Agent Prompt
### Issue description
For inverted metrics (`latency_ms`, `model_size_mb`), the tie branch (`mx == mn`) sets `normalized = 1.0` and then applies inversion, resulting in `0.0` for all jobs on that metric. This contradicts the tie comment and produces misleading scores; it can also affect winner selection when other metrics are missing.

### Issue Context
Latency/size are “lower is better” metrics. If all values are equal, every job should receive an equal *best/tie* contribution for that metric (not zero).

### Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/agent_compare.py[126-134]

### Suggested fix
Handle the `mx == mn` case after considering directionality, e.g. either:
- Skip inversion in the tie case, or
- Set `normalized = 0.0` for inverted metrics when `mx == mn` (so `1 - normalized` becomes `1.0`), and keep `normalized = 1.0` for non-inverted metrics.

Keep behavior consistent with the comment and ensure scores remain meaningful even when only some metrics are present.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@qodo-code-review

Copy link
Copy Markdown
Contributor

Qodo Fixer

✅ Committed (1) · ☑ Fixed (1)

Grey Divider

Commits pushed directly to this PR — no separate fix PR opened.

Process — 1 fixed
  • ☑ Fixed: Exit code 0 lost

@tonythethompson

Copy link
Copy Markdown
Owner Author

Superseded by consolidated PR #245

@tonythethompson
tonythethompson deleted the test/agent-model-info-unit-tests branch August 13, 2026 00:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant