Skip to content

feat: implement v0.4 release spec - #250

Merged
tonythethompson merged 3 commits into
mainfrom
codex/v04-release
Aug 11, 2026
Merged

tonythethompson merged 3 commits into
mainfrom
codex/v04-release

Conversation

@tonythethompson

@tonythethompson tonythethompson commented Aug 11, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • add Express-owned in-memory agent sessions and protected REST endpoints
  • connect planner, executor, and diagnosis MCP tools to session context
  • deduplicate GraphCanvas SVG definitions with shared markers and symbols
  • add release coverage for session behavior and SVG invariants
  • correct three stale cross-platform/accessibility test assertions

Verification

  • TypeScript: passed
  • ESLint: passed with 5 existing warnings
  • unit: 1,086 passed
  • server: 423 passed
  • integration: 65 passed
  • component: 144 passed
  • focused Python agent/session tests: passed
  • recipe validation: passed
  • production build: passed
  • no production references to olive-ai 0.12.1
  • git diff check: passed

Python environment note

The project .venv currently contains incompatible compiled NumPy/Pydantic wheels. A system-Python full sweep reached 607 passing tests, with unrelated semantic-worker timeout and Hypothesis too-slow failures. The v0.4-focused Python tests pass independently.

Review in cubic

- Remove outdated REVIEW.md document
- Update ROADMAP.md with current priorities and architectural themes
- Refactor ModelMemoryCompare, VramEstimateBanner, and BatchComparisonView for improved maintainability
- Update ExecutionWorkspace and related recipe graph inspectors (Conversion, Output, Provider, Pruning, Quantization)
- Enhance MCPDiagnosticCard with improved test coverage and diagnostic capabilities
- Update PerformanceMetrics and PassGuidanceCard for better metric visualization
- Refactor hardware integration panels (HardwarePassCards, HardwareProviderCard, IHVIntegrationPanel)
- Improve memory offload controls and input source panels
- Update playground components (ArenaConvenience, WebGpuBenchmarkPanel)
- Refactor environment and catalog browser panels for better UX
- Update logging metrics service and venv status detection
- Minor CSS improvements for component consistency

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @tonythethompson, you have reached your weekly rate limit of 500000 diff characters.

Please try again later or upgrade to continue using Sourcery

@vercel

vercel Bot commented Aug 11, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
olive-studio Ready Ready Preview Aug 11, 2026 8:59am

@coderabbitai

coderabbitai Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@tonythethompson, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 29 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 2296a5b2-5c5e-430c-840b-733b3b0797c8

📥 Commits

Reviewing files that changed from the base of the PR and between 7e332be and 995a8b9.

📒 Files selected for processing (43)
  • REVIEW.md
  • docs/ROADMAP.md
  • olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py
  • olive-mcp-server/olive_mcp_server/tools/agent_execute.py
  • olive-mcp-server/olive_mcp_server/tools/agent_planner.py
  • olive-mcp-server/olive_mcp_server/tools/studio_loopback.py
  • olive-mcp-server/tests/test_agent_execute.py
  • olive-mcp-server/tests/test_agent_sessions.py
  • src/components/features/ModelMemoryCompare.tsx
  • src/components/features/VramEstimateBanner.tsx
  • src/components/features/execute/BatchComparisonView.tsx
  • src/components/features/execute/BatchProcessingPanel.test.tsx
  • src/components/features/execute/ExecutionWorkspace.tsx
  • src/components/features/execute/MCPDiagnosticCard.test.tsx
  • src/components/features/execute/MCPDiagnosticCard.tsx
  • src/components/features/execute/PerformanceMetrics.tsx
  • src/components/features/execute/recipe-graph/GraphCanvas.test.tsx
  • src/components/features/execute/recipe-graph/GraphCanvas.tsx
  • src/components/features/execute/recipe-graph/PassGuidanceCard.tsx
  • src/components/features/execute/recipe-graph/inspectors/ConversionInspector.tsx
  • src/components/features/execute/recipe-graph/inspectors/OutputInspector.tsx
  • src/components/features/execute/recipe-graph/inspectors/ProviderInspector.tsx
  • src/components/features/execute/recipe-graph/inspectors/PruningInspector.tsx
  • src/components/features/execute/recipe-graph/inspectors/QuantizationInspector.tsx
  • src/components/features/execute/recipe-graph/svgDefs.ts
  • src/components/features/ihv/HardwarePassCards.tsx
  • src/components/features/ihv/HardwareProviderCard.tsx
  • src/components/features/ihv/IHVIntegrationPanel.tsx
  • src/components/features/ihv/MemoryOffloadControls.tsx
  • src/components/features/input/InputEnvironmentPanel.test.tsx
  • src/components/features/input/InputEnvironmentPanel.tsx
  • src/components/features/input/InputHuggingFaceSourceForm.tsx
  • src/components/features/input/RecipeCatalogBrowserPanels.tsx
  • src/components/features/playground/ArenaConvenience.tsx
  • src/components/features/playground/WebGpuBenchmarkPanel.tsx
  • src/index.css
  • src/lib/oliveLogMetrics.ts
  • src/server/__tests__/agentSessions.test.ts
  • src/server/__tests__/routes.integration.test.ts
  • src/server/routes/olive.ts
  • src/server/services/olive/agentSessions.ts
  • src/server/services/venv/pathIsolation.test.ts
  • src/server/services/venv/status.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8830371366

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread olive-mcp-server/olive_mcp_server/tools/agent_execute.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/agent_execute.py Outdated
Comment thread src/server/services/olive/agentSessions.ts
@qodo-code-review

Copy link
Copy Markdown
Contributor

PR Summary by Qodo

Implement v0.4 agent sessions, MCP session context, and SVG defs dedup

✨ Enhancement 🧪 Tests 📝 Documentation 🐞 Bug fix 🕐 40+ Minutes

Grey Divider

AI Description

• Add local-only agent sessions API with in-memory retry/diagnostic context.
• Wire MCP planner/executor/diagnosis tools to persist session metadata across calls.
• Deduplicate GraphCanvas SVG defs and harden UI/accessibility test expectations.
Diagram

graph TD
  MCP["MCP tools (Python)"] --> Loop["studio_loopback.py"] --> API(["Express Olive API"]) --> Store[("agentSessions (memory)")]
  UI["GraphCanvas (React)"] --> Defs["svgDefs.ts"]
  Tests["Tests (TS/Py)"] --> API --> Store
  Tests --> UI --> Defs

  subgraph Legend
    direction LR
    _mod["Module/UI"] ~~~ _svc(["Service/API"]) ~~~ _db[("In-memory store")]
  end
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Persist sessions (SQLite/Redis)
  • ➕ Survives Express restarts; better for long-running/remote use
  • ➕ Can support multi-user or background agents safely
  • ➖ Adds operational and migration complexity
  • ➖ Likely unnecessary under the loopback-only threat model
2. Store session context on jobs instead of sessions
  • ➕ Avoids a new store concept; relies on existing job registry/metadata
  • ➕ Naturally scopes context to executions
  • ➖ Harder to support planner-only or diagnosis-only calls that need shared context
  • ➖ Mixes retry history with job lifecycle, complicating cleanup semantics

Recommendation: Keep the current Express-owned in-memory session store for v0.4: it matches the loopback/local-only model, is easy to reason about, and the PR adds good contract tests. Revisit persistence only if sessions must survive server restarts or support non-local clients.

Files changed (41) +814 / -198

Enhancement (7) +317 / -44
agent_diagnosis.pyPersist diagnosis results into agent session context +29/-1

Persist diagnosis results into agent session context

• Adds optional session_id support and integrates with Studio loopback session creation/update. On completion, records last recipe/failure and appends a bounded diagnostic note trail.

olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py

agent_execute.pyRecord execution attempts into agent sessions +33/-6

Record execution attempts into agent sessions

• Adds optional session_id and a finish() hook to record attempt success/failure and notes via the Studio session API. Ensures submission/polling error returns also update session metadata when enabled.

olive-mcp-server/olive_mcp_server/tools/agent_execute.py

agent_planner.pyAttach planning steps to agent session notes +27/-2

Attach planning steps to agent session notes

• Adds optional session_id and updates the Studio session with a bounded diagnostic note when a plan is created. Uses environment-driven enablement when OLIVE_STUDIO_API_URL is set.

olive-mcp-server/olive_mcp_server/tools/agent_planner.py

studio_loopback.pyAdd session create/read/update helpers for Studio bridge +52/-1

Add session create/read/update helpers for Studio bridge

• Introduces _ensure_session(), _update_session(), and _record_attempt() helpers that call local-only Express endpoints. Uses URL quoting for safe path construction and validates sessionId presence on creation.

olive-mcp-server/olive_mcp_server/tools/studio_loopback.py

InputHuggingFaceSourceForm.tsxRemove trust-remote-code UI block and simplify token messaging +5/-33

Remove trust-remote-code UI block and simplify token messaging

• Removes the Trust Remote Code switch block from the HF source form and drops unused icon imports. Tweaks token status/help text and adjusts layout alignment.

src/components/features/input/InputHuggingFaceSourceForm.tsx

olive.tsExpose local-only REST endpoints for agent sessions +77/-1

Expose local-only REST endpoints for agent sessions

• Adds POST/GET/PUT routes under /api/olive/agent/sessions guarded by studioLocalOnly. PUT supports both attempt recording and metadata patching via a typed parseBody contract.

src/server/routes/olive.ts

agentSessions.tsAdd Express-owned in-memory agent session store +94/-0

Add Express-owned in-memory agent session store

• Implements a process-lifetime Map-backed session store with UUID IDs, attempt counting, last recipe/failure tracking, and capped diagnostic notes. Includes a test-only reset helper.

src/server/services/olive/agentSessions.ts

Bug fix (7) +34 / -28
BatchProcessingPanel.test.tsxFix thumbs-up accessible name assertion +1/-1

Fix thumbs-up accessible name assertion

• Updates the test to match the updated ARIA label punctuation used by MCP diagnostic feedback buttons.

src/components/features/execute/BatchProcessingPanel.test.tsx

MCPDiagnosticCard.test.tsxAlign MCP diagnostic feedback tests with new ARIA labels +10/-10

Align MCP diagnostic feedback tests with new ARIA labels

• Updates button-name expectations from em-dash phrasing to colon phrasing across multiple test cases.

src/components/features/execute/MCPDiagnosticCard.test.tsx

MCPDiagnosticCard.tsxUpdate feedback ARIA labels and fix tooltip copy +4/-4

Update feedback ARIA labels and fix tooltip copy

• Changes feedback button labels to use colon phrasing for better accessibility/name matching. Adjusts apply-fix tooltip text for clearer guidance when diagnostics are guidance-only.

src/components/features/execute/MCPDiagnosticCard.tsx

PerformanceMetrics.tsxMatch metrics placeholder sentinel to parser output +4/-4

Match metrics placeholder sentinel to parser output

• Updates filtering logic to treat '-' as the missing-value sentinel, aligning with the log metrics parser change.

src/components/features/execute/PerformanceMetrics.tsx

InputEnvironmentPanel.test.tsxFix trust-remote-code toggle role/name assertion +2/-2

Fix trust-remote-code toggle role/name assertion

• Updates test to match the UI control’s accessible role/name (checkbox and simplified label matching).

src/components/features/input/InputEnvironmentPanel.test.tsx

oliveLogMetrics.tsChange default metrics sentinel from em dash to hyphen +4/-4

Change default metrics sentinel from em dash to hyphen

• Updates parseOliveMetricsFromLogs to return '-' for missing metrics, aligning UI components on a consistent placeholder.

src/lib/oliveLogMetrics.ts

pathIsolation.test.tsHarden PATH isolation test for cross-platform env casing +9/-3

Harden PATH isolation test for cross-platform env casing

• Avoids relying on process.env PATH casing and replaces hard-coded /usr/bin with a synthetic inherited bin path to make the test portable across platforms.

src/server/services/venv/pathIsolation.test.ts

Refactor (21) +144 / -99
ModelMemoryCompare.tsxPolish memory comparison copy and punctuation +2/-2

Polish memory comparison copy and punctuation

• Adjusts explanatory UI copy to be more consistent and readable without changing logic.

src/components/features/ModelMemoryCompare.tsx

VramEstimateBanner.tsxClarify VRAM estimate messaging +2/-2

Clarify VRAM estimate messaging

• Tweaks banner copy to be more explicit about heuristics and hybrid offload behavior.

src/components/features/VramEstimateBanner.tsx

BatchComparisonView.tsxStandardize missing-value placeholders in batch comparison +2/-2

Standardize missing-value placeholders in batch comparison

• Replaces em-dash placeholders with hyphen placeholders for consistency across the UI.

src/components/features/execute/BatchComparisonView.tsx

ExecutionWorkspace.tsxPolish execution log/status copy for consistency +5/-5

Polish execution log/status copy for consistency

• Updates several user-facing log/status strings (punctuation and phrasing) without altering execution behavior.

src/components/features/execute/ExecutionWorkspace.tsx

GraphCanvas.tsxDeduplicate SVG wire defs and add shared marker/symbol usage +11/-8

Deduplicate SVG wire defs and add shared marker/symbol usage

• Moves gradient/marker/symbol definitions into a shared component and references them by exported IDs. Replaces animated circle with a reusable <use> symbol and adds arrow markers to paths.

src/components/features/execute/recipe-graph/GraphCanvas.tsx

PassGuidanceCard.tsxTighten guidance formatting +1/-1

Tighten guidance formatting

• Adjusts separator punctuation for guidance description rendering to improve readability.

src/components/features/execute/recipe-graph/PassGuidanceCard.tsx

ConversionInspector.tsxClarify conversion-skipped messaging +1/-1

Clarify conversion-skipped messaging

• Updates copy to use consistent punctuation in the skipped-state message.

src/components/features/execute/recipe-graph/inspectors/ConversionInspector.tsx

OutputInspector.tsxClarify heuristic metrics banner text +1/-1

Clarify heuristic metrics banner text

• Refines copy describing simulated heuristics vs profiled metrics.

src/components/features/execute/recipe-graph/inspectors/OutputInspector.tsx

ProviderInspector.tsxStandardize provider availability copy +1/-1

Standardize provider availability copy

• Changes unavailable provider label formatting for better readability in dropdowns.

src/components/features/execute/recipe-graph/inspectors/ProviderInspector.tsx

PruningInspector.tsxNormalize pruning preset/tooltip punctuation +12/-12

Normalize pruning preset/tooltip punctuation

• Updates preset descriptions, skipped-state copy, and tooltip phrasing to use colon-separated explanations consistently.

src/components/features/execute/recipe-graph/inspectors/PruningInspector.tsx

QuantizationInspector.tsxNormalize quantization preset labels and descriptions +50/-50

Normalize quantization preset labels and descriptions

• Updates many preset labels/options and explanatory text to consistent colon-separated phrasing and improves readability of informational sections.

src/components/features/execute/recipe-graph/inspectors/QuantizationInspector.tsx

svgDefs.tsIntroduce shared SVG defs component and exported IDs +42/-0

Introduce shared SVG defs component and exported IDs

• Adds GraphSvgDefs plus exported IDs for gradient, arrow marker, and dot symbol used across the graph wire rendering.

src/components/features/execute/recipe-graph/svgDefs.ts

HardwarePassCards.tsxImprove unsupported-pass messaging +1/-1

Improve unsupported-pass messaging

• Tweaks copy for unsupported hardware/pass combinations for clarity.

src/components/features/ihv/HardwarePassCards.tsx

HardwareProviderCard.tsxPolish provider capability copy +4/-4

Polish provider capability copy

• Updates several explanatory strings around legacy CUDA constraints and deploy-only targets for clearer messaging.

src/components/features/ihv/HardwareProviderCard.tsx

IHVIntegrationPanel.tsxClarify provider list status copy +1/-1

Clarify provider list status copy

• Refines detected/undetected provider messaging while keeping selection logic intact.

src/components/features/ihv/IHVIntegrationPanel.tsx

MemoryOffloadControls.tsxTidy CUDA version option copy +2/-2

Tidy CUDA version option copy

• Minor punctuation tweaks to CUDA 13.x option descriptions.

src/components/features/ihv/MemoryOffloadControls.tsx

InputEnvironmentPanel.tsxPolish applied-recipe hint copy +1/-1

Polish applied-recipe hint copy

• Minor copy change to improve readability of the applied recipe hint.

src/components/features/input/InputEnvironmentPanel.tsx

RecipeCatalogBrowserPanels.tsxClarify hardware probe unavailable copy +1/-1

Clarify hardware probe unavailable copy

• Adjusts a single message string for consistency.

src/components/features/input/RecipeCatalogBrowserPanels.tsx

ArenaConvenience.tsxPolish provider-apply status text +1/-1

Polish provider-apply status text

• Adjusts success status copy to use consistent sentence punctuation.

src/components/features/playground/ArenaConvenience.tsx

WebGpuBenchmarkPanel.tsxStandardize benchmark log phrasing +2/-2

Standardize benchmark log phrasing

• Tweaks log lines to use consistent punctuation while preserving behavior.

src/components/features/playground/WebGpuBenchmarkPanel.tsx

status.tsPolish runtime status hint copy +1/-1

Polish runtime status hint copy

• Minor text tweak to use consistent punctuation when the default runtime is missing.

src/server/services/venv/status.ts

Tests (4) +262 / -0
test_agent_sessions.pyAdd Python contract tests for session bridge helpers +67/-0

Add Python contract tests for session bridge helpers

• Adds unit tests for session creation when no ID is provided, reading an existing session, and ensuring attempt updates include the dispatch flag/body shape.

olive-mcp-server/tests/test_agent_sessions.py

GraphCanvas.test.tsxAdd SVG defs deduplication and interaction invariants test +50/-0

Add SVG defs deduplication and interaction invariants test

• Adds a component test ensuring a single <defs> block with unique IDs, correct marker/symbol usage, and preserved click/keyboard/resize interactions.

src/components/features/execute/recipe-graph/GraphCanvas.test.tsx

agentSessions.test.tsAdd unit tests for in-memory agent session store +71/-0

Add unit tests for in-memory agent session store

• Covers UUID uniqueness, missing-session behavior, metadata patching without attempt increments, and attempt recording with bounded note retention.

src/server/tests/agentSessions.test.ts

routes.integration.test.tsAdd integration coverage for agent session routes +74/-0

Add integration coverage for agent session routes

• Adds create/get/put flows for sessions, validating attempt recording vs metadata-only updates and error handling for unknown sessions and malformed bodies.

src/server/tests/routes.integration.test.ts

Documentation (1) +56 / -27
ROADMAP.mdRefresh roadmap for v0.3 shipped and v0.4 scope +56/-27

Refresh roadmap for v0.3 shipped and v0.4 scope

• Reformats and updates roadmap tables and phase notes, marking v0.3 items shipped and re-scoping upcoming agent UI milestones. Adds a clearer v0.4 validation/test hardening checklist and adjusts later-version headings.

docs/ROADMAP.md

Other (1) +1 / -0
index.cssSet global base font size in small-screen media query +1/-0

Set global base font size in small-screen media query

• Adds font-size: 18px to html under the existing responsive block, affecting global typography scale.

src/index.css

@greptile-apps

greptile-apps Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR implements the v0.4 agent-loop session contract and consolidates GraphCanvas SVG resources while updating related tests and release documentation.

  • Adds bounded, Express-owned agent sessions and connects planning, execution, and diagnosis tools to shared session state.
  • Persists bounded failure context and attempt history for agent retries and diagnosis.
  • Deduplicates recipe-graph SVG definitions through shared markers and symbols.
  • Updates release coverage and stale cross-platform and accessibility assertions.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
olive-mcp-server/olive_mcp_server/tools/agent_execute.py Records bounded failure details for failed, timed-out, and bridge-error outcomes while preserving the primary execution result if session bookkeeping fails.
olive-mcp-server/olive_mcp_server/tools/studio_loopback.py Adds bridge helpers to create, retrieve, update, and record attempts against Express-owned sessions.
olive-mcp-server/olive_mcp_server/tools/agent_planner.py Associates planning calls with session context and records a bounded diagnostic note after successful planning.
olive-mcp-server/olive_mcp_server/tools/agent_diagnosis.py Persists the diagnosed failure, resulting recipe, and diagnostic note in the active agent session.
src/server/services/olive/agentSessions.ts Implements in-memory session state with a 200-entry capacity, 24-hour idle expiry, bounded histories, and oldest-session eviction.
src/server/routes/olive.ts Exposes protected session creation, retrieval, and update endpoints with request validation.
src/components/features/execute/recipe-graph/GraphCanvas.tsx Reuses shared SVG definitions to avoid duplicating markers and symbols across graph elements.
src/components/features/execute/recipe-graph/svgDefs.ts Centralizes the marker and symbol definitions consumed by GraphCanvas.

Reviews (2): Last reviewed commit: "fix agent execution context and bound se..." | Re-trigger Greptile

Comment thread olive-mcp-server/olive_mcp_server/tools/agent_execute.py Outdated
Comment thread src/server/services/olive/agentSessions.ts
@qodo-code-review

qodo-code-review Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Failed runs erase context ✓ Resolved 🐞 Bug ≡ Correctness
Description
finish() only derives failure context from a top-level message/error or poll_error, while
normal terminal failures expose status, exit_code, and logs instead, so failures can be
recorded with no diagnostic text. Then recordAttempt() sets lastFailure to null whenever
failure is absent, even when success: false, erasing the very retry context needed for
multi-step agent runs.
Code

olive-mcp-server/olive_mcp_server/tools/agent_execute.py[R239-240]

+            failure = result.get("message") if result.get("error") else result.get("poll_error")
+            success = result.get("status") == "completed" and not result.get("error")
Relevance

●●● Strong

Missing failure context/erased retry state is a clear correctness bug; similar “distinguish failure
vs empty” fixes were accepted.

PR-#75

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The evidence indicates that _observation_result() represents failed jobs using fields like
status, exit_code, and logs rather than a top-level message/error, and the existing
failed-job test demonstrates that result shape; as a consequence, finish() can pass along no
failure string when neither message nor poll_error is present. In the session store,
agentSessions.recordAttempt assigns lastFailure: data.failure ?? null irrespective of the
success flag, which means a failed attempt recorded without explicit failure text will clear any
previously stored failure detail, reducing the usefulness of diagnosticNotes/lastFailure across
retries.

olive-mcp-server/olive_mcp_server/tools/agent_execute.py[187-207]
olive-mcp-server/tests/test_agent_execute.py[80-99]
src/server/services/olive/agentSessions.ts[76-84]
src/server/services/olive/agentSessions.ts[76-86]
olive-mcp-server/olive_mcp_server/tools/agent_execute.py[236-251]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Normal failed/cancelled/timed-out executions can end up recording `lastFailure` as `null` because `finish()` only inspects top-level `message`/`error` or `poll_error`, while actual terminal failure results may instead provide `status`, `exit_code`, and `logs`; then `recordAttempt()` unconditionally replaces `lastFailure` with `data.failure ?? null` even when `success: false`, wiping prior failure context on failed attempts that don’t supply failure text.

## Issue Context
Update failure extraction so that when execution is not successful you still derive useful failure context from non-success `status`, timeout/cancel conditions, exit code, and a bounded tail of logs. In the session store, avoid clearing the current `lastFailure` merely because the incoming attempt lacks a top-level failure string, especially when recording a failed attempt; callers like `execute_and_observe`’s `finish()` can legitimately record `success=false` with `failure=None` when neither `message` nor `poll_error` exists.

## Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/agent_execute.py[187-207]
- olive-mcp-server/olive_mcp_server/tools/agent_execute.py[236-247]
- src/server/services/olive/agentSessions.ts[69-84]
- src/server/services/olive/agentSessions.ts[76-86]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Bookkeeping masks execution result ✓ Resolved 🐞 Bug ≡ Correctness
Description
execute_and_observe.finish() replaces an already-submitted job's result with the follow-up session
PUT error when bookkeeping fails. This discards job_id, status, and side_effect: true, so an
agent can retry an optimization whose actual outcome was merely hidden.
Code

olive-mcp-server/olive_mcp_server/tools/agent_execute.py[R248-249]

+            if isinstance(update.get("error"), str) and update["error"]:
+                return update
Relevance

●●● Strong

Preventing bookkeeping errors from masking real execution outcomes matches prior accepted “don’t let
secondary errors hide primary result” patterns.

PR-#75
PR-#135

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The execution result is constructed with job_id and side_effect: True, then finish() performs
a separate session PUT and currently returns that PUT error instead. Because the PUT occurs after
submission/observation, its failure does not mean the Olive job failed or did not run.

olive-mcp-server/olive_mcp_server/tools/agent_execute.py[187-207]
olive-mcp-server/olive_mcp_server/tools/agent_execute.py[236-251]
olive-mcp-server/olive_mcp_server/tools/studio_loopback.py[292-298]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Session bookkeeping failures currently replace the execution result after Olive may already have run, hiding the job identity and side effect from callers.

## Issue Context
Treat the session update as secondary metadata persistence. Preserve the original execution outcome and expose any update failure as non-fatal metadata so callers do not retry an uncertain side effect.

## Fix Focus Areas
- olive-mcp-server/olive_mcp_server/tools/agent_execute.py[236-251]
- olive-mcp-server/olive_mcp_server/tools/studio_loopback.py[292-298]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

3. PUT session route silently drops metadata fields ✗ Dismissed 🐞 Bug ≡ Correctness
Description
In the PUT /olive/agent/sessions/:sessionId handler, when attempt is true the handler calls
recordAttempt with only recipe, failure, success, and note, ignoring the patch object
built from lastRecipe, lastFailure, and diagnosticNotes. A caller sending attempt: true
together with those metadata fields has them silently discarded instead of applied or rejected,
causing agent-loop session state to diverge from what was requested.
Code

src/server/routes/olive.ts[R314-327]

+    const { attempt, recipe, failure, note } = body.parsed;
+    const patch = {
+      ...(body.parsed.lastRecipe !== undefined ? { lastRecipe: body.parsed.lastRecipe } : {}),
+      ...(body.parsed.lastFailure !== undefined ? { lastFailure: body.parsed.lastFailure } : {}),
+      ...(body.parsed.success !== undefined ? { success: body.parsed.success } : {}),
+      ...(body.parsed.diagnosticNotes !== undefined
+        ? { diagnosticNotes: body.parsed.diagnosticNotes }
+        : {}),
+    };
+    const session = !sessionId
+      ? undefined
+      : attempt
+        ? recordAttempt(sessionId, { recipe, failure, success: body.parsed.success, note })
+        : updateSession(sessionId, patch);
Relevance

●●● Strong

Silently discarding provided fields in a route is a straightforward correctness bug; team has
accepted route contract fixes.

PR-#135

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The route builds patch from lastRecipe/lastFailure/success/diagnosticNotes but only passes it to
updateSession() in the non-attempt branch; the attempt branch calls recordAttempt() with a
different, smaller field set, so lastRecipe/lastFailure/diagnosticNotes supplied alongside
attempt:true are never applied.

src/server/routes/olive.ts[314-327]
src/server/services/olive/agentSessions.ts[69-89]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The PUT `/olive/agent/sessions/:sessionId` handler silently ignores `lastRecipe`, `lastFailure`, and `diagnosticNotes` fields when `attempt: true` is also present in the body, because `recordAttempt` is called with a different field set than the `patch` object built for the metadata-only path.

## Issue Context
The route branches on `body.parsed.attempt`: true routes to `recordAttempt(sessionId, { recipe, failure, success, note })`, false routes to `updateSession(sessionId, patch)`. The `patch` variable (containing lastRecipe/lastFailure/diagnosticNotes) is computed unconditionally but only used in the `updateSession` branch.

## Fix Focus Areas
- src/server/routes/olive.ts[314-327]
- src/server/services/olive/agentSessions.ts[69-89]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


4. Roadmap reverses trust policy ✗ Dismissed 🐞 Bug ⚙ Maintainability
Description
The roadmap now states that Hugging Face recipes automatically emit trust_remote_code: true, but
production defaults the setting to false and emits it only after explicit user opt-in. This
security-sensitive release record documents the opposite trust boundary and can drive incorrect
validation or future implementation changes.
Code

docs/ROADMAP.md[54]

+- [x] `trust_remote_code` default-flip handled: auto-emit `true` for HuggingFace models + info advisory
Relevance

●●● Strong

Team has precedent updating ROADMAP bullets to match real implementation details/security posture.

PR-#220

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The changed roadmap claims automatic emission, while the builder checks trustRemoteCode === true,
the default is false, and the UI presents an explicit warning checkbox.

docs/ROADMAP.md[50-55]
src/lib/oliveRecipeBuilder.ts[733-744]
src/lib/defaultPasses.ts[21-30]
src/components/features/input/InputHuggingFaceSourceForm.tsx[68-84]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The release roadmap says `trust_remote_code` is auto-enabled even though the implemented policy requires explicit user opt-in.

## Issue Context
Correct the release statement to describe the false default, explicit Hugging Face source control, and conditional recipe emission. Do not change runtime behavior to match the unsafe wording.

## Fix Focus Areas
- docs/ROADMAP.md[50-55]
- src/lib/oliveRecipeBuilder.ts[733-744]
- src/lib/defaultPasses.ts[21-30]
- src/components/features/input/InputHuggingFaceSourceForm.tsx[68-84]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


5. Agent sessions grow unbounded ✓ Resolved 🐞 Bug ☼ Reliability
Description
With the Studio bridge configured, tool calls that omit session_id (and the unrate-limited POST
/olive/agent/sessions endpoint) can create sessions indefinitely in a process-lifetime in-memory
Map that has no TTL, capacity bound, or eviction/cleanup path. In long-running processes this
retains every session’s recipe/notes indefinitely and planner errors after session creation can
leave entries whose IDs are never returned, risking gradual memory exhaustion (e.g., via a local
script or MCP client bug).
Code

src/server/services/olive/agentSessions.ts[R25-26]

+const MAX_DIAGNOSTIC_NOTES = 50;
+const sessions = new Map<string, AgentSession>();
Relevance

●● Moderate

Rate-limit/TTL for in-memory sessions is plausible, but involves policy/scope tradeoffs beyond a
trivial fix.

PR-#108

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The cited code paths show that sessions are stored in a plain in-memory Map that only supports
insert/read/update and provides no production deletion, TTL, cap, or eviction behavior, so entries
persist for the entire Express process lifetime. The Studio bridge/planner flow auto-creates a
session when no session_id is provided, and because session creation happens before later
parse/strategy errors, some sessions can be created and then become unreachable when an error
returns before the ID is propagated back to the caller; additionally, session creation is reachable
from POST /olive/agent/sessions without a rate limiter, unlike other mutation-heavy routes that
use heavyCommandRateLimit or similar, making it possible to accumulate sessions indefinitely. The
repository’s existing job registry is referenced as demonstrating a TTL-sweeper pattern for bounding
process-lifetime state.

olive-mcp-server/olive_mcp_server/tools/studio_loopback.py[250-280]
olive-mcp-server/olive_mcp_server/tools/agent_planner.py[520-536]
src/server/services/olive/agentSessions.ts[25-45]
src/server/services/olive/state.ts[24-31]
src/server/services/olive/agentSessions.ts[25-42]
src/server/routes/olive.ts[264-271]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The in-memory agent session store grows without bound because sessions are held in a process-lifetime `Map` with no TTL, maximum size, or eviction/production cleanup path, and sessions can be created indefinitely (including via POST `/olive/agent/sessions` without a rate limiter and via Studio bridge tool calls that omit `session_id`). This can lead to long-running processes retaining every session’s recipe/notes indefinitely, and planner errors after session creation can leave sessions whose IDs are never returned, creating unreachable entries and increasing the risk of memory exhaustion.

## Issue Context
Sessions are intentionally stored for the lifetime of the Express process to survive MCP stdio restarts, but nothing currently bounds how many sessions can accumulate. Add lifecycle management such as TTL eviction plus a bounded maximum size, and ensure auto-created sessions are either returned to callers or cleaned up when later tool/planner errors occur. Keep test reset behavior separate from production cleanup, and consider aligning session-creation endpoints with existing rate-limiting patterns (e.g., `heavyCommandRateLimit`) used for other mutation-heavy routes.

## Fix Focus Areas
- src/server/services/olive/agentSessions.ts[25-45]
- src/server/services/olive/agentSessions.ts[91-94]
- src/server/routes/olive.ts[264-271]
- olive-mcp-server/olive_mcp_server/tools/studio_loopback.py[250-280]
- olive-mcp-server/olive_mcp_server/tools/agent_planner.py[520-536]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context used
✅ Compliance rules (platform): 29 rules
Review mode: 🧠 Deep: This release changes protected REST session APIs, MCP session propagation, job-attempt state, and SVG rendering across multiple independent code paths, creating a dense set of subtle correctness and security risks.

Grey Divider

Tip of the day
💡 Did you know, you can group findings by type and pick your Finding display, from Minimal to Full

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread olive-mcp-server/olive_mcp_server/tools/agent_execute.py Outdated
Comment thread olive-mcp-server/olive_mcp_server/tools/agent_execute.py Outdated
Comment thread src/server/services/olive/agentSessions.ts
Comment thread docs/ROADMAP.md
Comment thread src/server/routes/olive.ts
@qodo-code-review

Copy link
Copy Markdown
Contributor

Qodo Fixer

✅ Merged (0) · ☑ Fixed (0)

Process

  • No fixes were applied (no_fixes_applied)

@tonythethompson
tonythethompson merged commit 1b2e2ef into main Aug 11, 2026
13 checks passed
@tonythethompson
tonythethompson deleted the codex/v04-release branch August 11, 2026 09:52
@linear-code

linear-code Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

OLI-87

This branch was successfully deployed

1 active deployment
Preview — 995a8b9d Deployed Aug 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant