Skip to content

feat: add Execute agent mode, activity log, and batch comparison UI - #285

Merged
tonythethompson merged 129 commits into
mainfrom
feat/execute-agent-mode-batch-comparison
Aug 14, 2026
Merged

tonythethompson merged 129 commits into
mainfrom
feat/execute-agent-mode-batch-comparison

Conversation

@tonythethompson

@tonythethompson tonythethompson commented Aug 13, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Add ActivityLog + ActivityLogEntry: render agent stream entries with floor-based elapsed formatting
  • Add AgentControls + useAgentMode: start/stop agent execution with mode toggle and confirmation
  • Add AgentConfirmDialog: accessible confirmation dialog with test IDs and aria-describedby
  • Add useAgentStream: EventSource hook with exponential backoff reconnection (1s, 2s, 4s) and 50ms grace period for done-event race handling
  • Add BatchComparisonView: compare 2-10 job records with sortable columns, duration deltas, scoring preference, winner highlighting, excluded jobs, and accessible disabled-control messaging
  • Add ExportReportMenu: report export with test IDs
  • Add ModeToggle: agent mode toggle with test ID
  • Add CatalogUpdateNotice: catalog staleness notice
  • Update ExecutionWorkspace to integrate new components

Stacked on PR #282 (shared validation libs) which provides the underlying batch comparison, activity log, and agent type libraries.

Test plan

  • pnpm test:component - 239 component tests pass (32 test files), including BatchComparisonView (19), AgentConfirmDialog (10), ModeToggle (8), useAgentMode (15), useAgentStream (19)
  • pnpm lint - 0 errors
  • tsc --noEmit - 0 errors

Generated with Devin

Review in cubic

Note

Add Execute agent mode with activity log, batch comparison UI, and export report menu

  • Adds an Agent mode to ExecutionWorkspace where users can start/stop an autonomous agent session, stream real-time activity entries via SSE (useAgentStream), and switch back to manual mode with a confirmation dialog if the agent is running.
  • Introduces useAgentMode hook to manage agent session state (running flag, bounded activity log, outcome, timestamps, job ID) with a 10s start confirmation timeout.
  • Adds ActivityLog, ActivityLogEntry, AgentControls, AgentConfirmDialog, and ModeToggle components to support the agent mode UI.
  • Extends BatchComparisonView and BatchProcessingPanel to support MCP compare_results calls with a scoring preference selector, winner highlighting, and eligibility enforcement (2–10 completed jobs); BatchJob now records startedAtMs/finishedAtMs timestamps.
  • Adds ExportReportMenu to the workspace header (behind a reportExport feature flag) that downloads Markdown or triggers print-as-PDF for completed job records.
  • Behavioral Change: In manual mode, the completed-state auxiliary action now navigates to Playground instead of opening the OWR export overlay.

Macroscope summarized 715cfe1.

Add core libraries for the v0.5 release:

- src/lib/types: agent and finding type definitions
- src/lib/featureFlags: safe boolean parsing with URL/localStorage/env precedence
- src/lib/batchComparison: job count validation (2-10), finite number checks,
  MCP output structural validation
- src/lib/multiLoraValidation: finite VRAM handling, adapter name/path/alpha
  validation, count limits
- src/lib/activityLog: floor-based elapsed formatting, optional stepRef
- src/lib/reportGenerator: correct default title, UTC timestamps, PDF export
- src/lib/reviewReconciler: deterministic findings authoritative, AI findings
  targeting deterministic pass fields suppressed
- src/lib/workspaceFingerprint: deterministic fingerprints excluding transient
  fields, stale async result handling via computation ID ref
- src/lib/recipeCatalogPin: GitHub rate-limit handling, truncated tree rejection,
  content decoding hardening, stale detection with upstream SHA
- src/lib/oliveRecipeBuilder: MultiLoRA gating behind PEFT + valid adapters,
  safe vramEstimateGb default
- src/lib/actionExecutor: state mutations routed through commitUiStateUpdate
- src/lib/hooks: useCatalogPin (refresh guard, resolve debounce, SHA propagation)
  and useWorkspaceFingerprint (memoization with stale result rejection)

All libraries include unit and property-based tests.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@vercel

vercel Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
olive-studio Ready Ready Preview Aug 13, 2026 11:04pm

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @tonythethompson, your pull request is larger than the review limit of 150000 diff characters

@qodo-code-review

Copy link
Copy Markdown
Contributor

ⓘ Your Qodo trial ends soon. Ask your workspace admin to set up billing to keep reviews running after the trial. Manage billing

@coderabbitai

coderabbitai Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@tonythethompson, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 62 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 87f904e7-f1b8-4520-8a51-8042b2f1d2f5

📥 Commits

Reviewing files that changed from the base of the PR and between 89cc67c and f7663e8.

📒 Files selected for processing (15)
  • package.json
  • src/components/features/execute/ActivityLog.test.tsx
  • src/components/features/execute/BatchComparisonView.tsx
  • src/components/features/execute/BatchProcessingPanel.tsx
  • src/components/features/execute/ExecutionWorkspace.test.tsx
  • src/components/features/execute/ExecutionWorkspace.tsx
  • src/components/features/execute/ExportReportMenu.tsx
  • src/components/features/execute/useAgentMode.test.ts
  • src/components/features/execute/useAgentMode.ts
  • src/components/features/execute/useAgentStream.test.ts
  • src/components/features/execute/useAgentStream.ts
  • src/lib/__tests__/featureFlagGating.test.ts
  • src/lib/oliveJobLogLimits.ts
  • src/server/services/olive/gpu.test.ts
  • src/server/services/olive/gpu.ts
📝 Walkthrough

Walkthrough

The execution workspace now supports Manual and Agent modes, agent lifecycle and SSE activity streaming, bounded activity logs, confirmation dialogs, report export, and batch-result comparison. New catalog pinning, workspace fingerprinting, shared finding contracts, and property-based validation tests are also included.

Changes

Execution features

Layer / File(s) Summary
Agent activity streaming
src/components/features/execute/useAgentStream.ts, src/components/features/execute/useAgentStream.test.ts
The SSE hook validates payloads, adapts log and metrics events, suppresses duplicates, handles terminal events, retries failed connections, and cleans up resources.
Agent lifecycle and bounded activity state
src/components/features/execute/useAgentMode.ts, src/components/features/execute/useAgentMode.test.ts, src/lib/activityLog.ts
The agent hook manages startup, confirmation timeouts, cancellation, stale sessions, outcomes, timestamps, and bounded FIFO activity logs.
Agent controls and activity presentation
src/components/features/execute/ActivityLogEntry.tsx, src/components/features/execute/ActivityLog.tsx, src/components/features/execute/AgentControls.tsx, src/components/features/execute/ModeToggle.tsx, src/components/features/execute/AgentConfirmDialog.tsx, src/components/features/execute/*test.tsx
New components render activity entries, status controls, mode selection, and an accessible focus-managed confirmation dialog.
Batch lifecycle and comparison orchestration
src/types.ts, src/components/features/execute/BatchProcessingPanel.tsx, src/components/features/execute/BatchProcessingPanel.test.tsx
Batch jobs record start and finish timestamps. The panel maps terminal jobs to history records and calls compare_results for completed selections.
Comparison result presentation
src/components/features/execute/BatchComparisonView.tsx, src/components/features/execute/BatchComparisonView.test.tsx, src/lib/__tests__/batchComparison.test.ts
The comparison view adds scoring controls, eligibility validation, MCP metrics, excluded-job details, winner highlighting, and null-value handling.
Workspace mode and export integration
src/components/features/execute/ExecutionWorkspace.tsx, src/components/features/execute/ExportReportMenu.tsx, src/components/features/execute/ExecutionWorkspace.test.tsx
The workspace wires Agent mode, streaming, controls, activity logs, confirmation handling, manual-mode boundaries, and feature-flagged report export.

Workspace infrastructure

Layer / File(s) Summary
Finding and action contracts
src/lib/types/findingTypes.ts, src/lib/__tests__/findingContract.test.ts
Shared contracts define findings, review results, action payloads, workspace fingerprint state, and excluded UI-state keys. Tests cover field constraints, patch sanitization, coercion, and state-preserving actions.
Workspace fingerprint computation
src/lib/workspaceFingerprint.ts, src/lib/hooks/useWorkspaceFingerprint.ts, src/lib/__tests__/workspaceFingerprint.test.ts
Fingerprint utilities serialize relevant state deterministically and compute SHA-256 values. The hook debounces computation and reports stale results.
Catalog pinning and freshness state
src/lib/hooks/useCatalogPin.ts, src/components/features/input/CatalogUpdateNotice.tsx, src/lib/__tests__/recipeCatalogPin.test.ts
Catalog metadata is persisted and validated. The hook resolves upstream revisions, detects staleness, and refreshes pinned catalog entries. The notice renders update and refresh states.
Validation and report coverage
src/lib/__tests__/multiLoraValidation.test.ts, src/lib/__tests__/reportGenerator.test.ts, src/lib/__tests__/reviewReconciler.test.ts
Property-based tests cover MultiLoRA validation, report generation, reconciliation behavior, immutability, ordering, and severity handling.

Estimated code review effort: 5 (Critical) | ~120 minutes

Mergeability Score: 🟡 Moderate · up to 89cc6

This PR adds agent execution and comparison/reporting flows, but the current head can duplicate activity entries after reconnects, show stale comparison winners after the job set changes, and leave comparison state stuck when response parsing stalls; dialog semantics and inconsistent report-export behavior also remain unresolved. These can produce misleading results or inaccessible controls, so merge should wait for fixes or explicit owner acceptance.

Possibly related PRs

Suggested reviewers: greptile-apps

🚥 Pre-merge checks | ✅ 8
✅ Passed checks (8 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed Docstring coverage is 75.27% which is sufficient. The required threshold is 60.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Pipeline Stage Enum Ordering ✅ Passed The full PR diff and both endpoint trees contain no SessionWorkflowStage enum, member references, or related comparisons; the ordering and legacy-mapping checks are not applicable.
Gpu/Cpu Runtime Boundary ✅ Passed The feature range changes only Studio TypeScript/tests and related files; no inference/ files, managed requirements files, main.py, or C# Diarization code were modified.
Managed Host Restart Safety ✅ Passed The PR changes only Execute UI, agent session, SSE, comparison, and fingerprint code; no managed-host classes, lease trackers, health busy fields, or host restart methods exist or are modified.
Description check ✅ Passed The description directly summarizes the agent mode, activity logging, batch comparison, export, workspace integration, and validation changes.
Title check ✅ Passed The title clearly summarizes the primary changes: Execute agent mode, activity logging, and batch comparison UI.
✨ Finishing Touches 💡 1
⚔️ Resolve merge conflicts 💡
  • Resolve merge conflict in branch feat/execute-agent-mode-batch-comparison
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/execute-agent-mode-batch-comparison
✨ Simplify code
  • Create PR with simplified code
  • Commit simplified code in branch feat/execute-agent-mode-batch-comparison

Warning

Review ran into problems

🔥 Problems

Linked repositories: Public OSS repositories can only analyze public repositories installed in this organization. Analyzed tonythethompson/QuickShell, tonythethompson/numan, tonythethompson/dependency-chain-substrate, skipped Trackdubllc/Trackdub.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@qodo-code-review

Copy link
Copy Markdown
Contributor

PR Summary by Qodo

Add Agent execute mode with SSE activity log and enhanced batch comparison UI

✨ Enhancement 🧪 Tests 🕐 40+ Minutes

Grey Divider

AI Description

• Add Agent execution mode with start/stop controls and safe manual-mode switch confirmation.
• Stream agent activity via SSE with bounded log buffering and resilient reconnection.
• Enhance batch comparison with scoring preference, winner highlighting, and exclusions display.
Diagram

graph TD
  EW["ExecutionWorkspace"] --> MT["ModeToggle"] --> AM["useAgentMode"] --> AL["ActivityLog"]
  EW --> AC["AgentControls"]
  EW --> CD["AgentConfirmDialog"]
  EW --> AS["useAgentStream"] --> SSE{{"SSE /api/olive/agent/stream"}}
  EW --> BC["BatchComparisonView"]
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Adopt a resilient SSE helper/library
  • ➕ Centralizes reconnect/backoff/done-event race handling for reuse across the app
  • ➕ Potentially reduces bespoke edge-case code and test surface
  • ➖ Adds dependency/abstraction overhead for a single stream use case
  • ➖ May be harder to fit custom 'done' event + grace-period semantics cleanly
2. Move agent session state into a global store (e.g., pipelineStore)
  • ➕ Persists agent mode/running state across navigation or workspace remounts
  • ➕ Enables other panels to reflect agent status consistently
  • ➖ Increases coupling and risk of cross-feature regressions
  • ➖ Current approach explicitly keeps agent state local, which is simpler and safer initially
3. Use WebSocket instead of SSE for agent stream
  • ➕ Bidirectional channel could support acknowledgements/commands in same transport
  • ➕ May be more robust for long-running interactive sessions
  • ➖ More infra complexity than needed for one-way streaming log events
  • ➖ Harder to debug and may require additional server changes

Recommendation: Current approach (local agent state + SSE hook with bounded retries and a small grace period) is a good fit for a one-way activity stream and keeps state changes scoped to the Execute UI. Consider extracting the SSE reconnection logic into a shared helper only if more streams appear or if other pages need the same done/race semantics.

Files changed (18) +3593 / -503

Enhancement (11) +2081 / -500
ActivityLog.tsxAdd scrollable ActivityLog with conditional auto-scroll +102/-0

Add scrollable ActivityLog with conditional auto-scroll

• Implements an agent activity log container with empty state, max-height scrolling, and auto-scroll-to-bottom only when the user remains near the bottom. Adds accessible log semantics via role="log" and aria-live.

src/components/features/execute/ActivityLog.tsx

ActivityLogEntry.tsxAdd ActivityLogEntry renderer with kind-specific styling and truncation toggle +85/-0

Add ActivityLogEntry renderer with kind-specific styling and truncation toggle

• Renders individual activity entries with timestamp, icon/color by kind, optional stepRef display, and an expand/collapse control when expandedText is available.

src/components/features/execute/ActivityLogEntry.tsx

AgentConfirmDialog.tsxAdd confirmation dialog for leaving agent mode while running +124/-0

Add confirmation dialog for leaving agent mode while running

• Implements an accessible modal confirmation dialog (aria-labelledby/aria-describedby) with basic focus management and Escape/backdrop dismissal, exposing onConfirm/onCancel callbacks.

src/components/features/execute/AgentConfirmDialog.tsx

AgentControls.tsxAdd AgentControls start/stop UI with status indicator +135/-0

Add AgentControls start/stop UI with status indicator

• Provides start/stop buttons and a status dot/label that reflects idle/running/success/failure/cancelled states derived from agentRunning and outcome.

src/components/features/execute/AgentControls.tsx

BatchComparisonView.tsxEnhance batch comparison UI with scoring preference and winner/exclusions +238/-10

Enhance batch comparison UI with scoring preference and winner/exclusions

• Expands supported comparison size to 2–10 jobs, adds scoring preference selection and a compare trigger with accessible disabled messaging, and renders MCP compare_results metrics with winner highlighting and excluded job details.

src/components/features/execute/BatchComparisonView.tsx

ExecutionWorkspace.tsxIntegrate agent mode, confirmation dialog, and activity log into ExecutionWorkspace +617/-490

Integrate agent mode, confirmation dialog, and activity log into ExecutionWorkspace

• Adds agent-mode state management (useAgentMode), SSE streaming (useAgentStream), a Manual/Agent toggle with guarded switching, and an AgentExecution card with controls + activity log. Also reorganizes manual-mode UI so overlays and controls are hidden while in agent mode.

src/components/features/execute/ExecutionWorkspace.tsx

ExportReportMenu.tsxAdd ExportReportMenu with feature-flag gating and test IDs +140/-0

Add ExportReportMenu with feature-flag gating and test IDs

• Implements a dropdown export menu (markdown download / PDF print) gated behind a feature flag, disabled when no records exist, and instrumented with data-testid hooks.

src/components/features/execute/ExportReportMenu.tsx

ModeToggle.tsxAdd segmented ModeToggle for manual vs agent execution mode +81/-0

Add segmented ModeToggle for manual vs agent execution mode

• Adds a small segmented-button control implemented as a radiogroup with two radio-like buttons, including disabled-state styling and a stable test id.

src/components/features/execute/ModeToggle.tsx

useAgentMode.tsAdd useAgentMode hook for agent session lifecycle and activity log state +216/-0

Add useAgentMode hook for agent session lifecycle and activity log state

• Implements local agent session state: manual/agent mode, running flag, startedAt, terminal outcome, and FIFO-bounded log entries with truncation. Includes a 10s start timeout that emits an error entry when no first event arrives, plus stop/complete helpers that append terminal entries.

src/components/features/execute/useAgentMode.ts

useAgentStream.tsAdd useAgentStream SSE hook with exponential backoff and done-event handling +273/-0

Add useAgentStream SSE hook with exponential backoff and done-event handling

• Introduces an EventSource-based hook that parses agent stream events into ActivityLogEntry and delivers them via callbacks. Implements reconnect with 1s/2s/4s backoff (max 3 retries) and a short grace period to avoid reconnecting during a done-event race.

src/components/features/execute/useAgentStream.ts

CatalogUpdateNotice.tsxAdd CatalogUpdateNotice for stale catalog refresh prompting +70/-0

Add CatalogUpdateNotice for stale catalog refresh prompting

• Adds an inline status notice shown when the recipe catalog is stale, including optional SHA display and a refresh action with loading indicator and accessible status messaging.

src/components/features/input/CatalogUpdateNotice.tsx

Tests (7) +1512 / -3
ActivityLog.test.tsxAdd component tests for ActivityLog auto-scroll and empty state +138/-0

Add component tests for ActivityLog auto-scroll and empty state

• Introduces tests for empty-state rendering, accessibility role/labels, className passthrough, and auto-scroll behavior when new entries arrive (including when user is scrolled away).

src/components/features/execute/ActivityLog.test.tsx

AgentConfirmDialog.test.tsxAdd tests for AgentConfirmDialog accessibility and interactions +102/-0

Add tests for AgentConfirmDialog accessibility and interactions

• Covers open/closed rendering, aria attributes, confirm/cancel button behavior, backdrop dismissal, and Escape-key dismissal.

src/components/features/execute/AgentConfirmDialog.test.tsx

AgentControls.test.tsxAdd tests for AgentControls button states and status display +126/-0

Add tests for AgentControls button states and status display

• Verifies Start/Stop enablement based on running state, status indicator mapping for outcomes, click handler wiring, and accessibility labels.

src/components/features/execute/AgentControls.test.tsx

BatchComparisonView.test.tsxExpand BatchComparisonView tests for scoring and compare_results output +208/-3

Expand BatchComparisonView tests for scoring and compare_results output

• Extends coverage for scoring preference selection, compare button enable/disable logic (2–10 jobs), winner highlighting, and excluded-jobs/no-winner messaging when compare_results output is provided.

src/components/features/execute/BatchComparisonView.test.tsx

ModeToggle.test.tsxAdd ModeToggle component tests +70/-0

Add ModeToggle component tests

• Tests manual/agent selection behavior, disabled state, and accessibility roles/aria-checked semantics.

src/components/features/execute/ModeToggle.test.tsx

useAgentMode.test.tsAdd unit tests for useAgentMode lifecycle and log buffering +340/-0

Add unit tests for useAgentMode lifecycle and log buffering

• Introduces extensive unit tests covering default state, mode switching, start timeout behavior, FIFO log limits, outcome/terminal entry handling, and cancellation/completion flows.

src/components/features/execute/useAgentMode.test.ts

useAgentStream.test.tsAdd unit tests for SSE parsing and reconnection behavior +528/-0

Add unit tests for SSE parsing and reconnection behavior

• Adds comprehensive tests for EventSource lifecycle, JSON parsing into ActivityLogEntry objects, done-event completion, exponential backoff reconnection, and error surfacing after retry exhaustion.

src/components/features/execute/useAgentStream.test.ts

Comment thread src/components/features/execute/AgentConfirmDialog.tsx
Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
Comment thread src/components/features/execute/AgentConfirmDialog.tsx
Comment thread src/components/features/execute/useAgentStream.ts Outdated
Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
Comment thread src/components/features/execute/useAgentMode.ts
Comment thread src/components/features/execute/ActivityLog.tsx
Comment thread src/components/features/execute/ExecutionWorkspace.tsx
Comment thread src/components/features/execute/useAgentStream.ts Outdated
@greptile-apps

greptile-apps Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds Execute agent mode, streamed activity, batch-run comparison, report export, and associated workspace controls.

  • Integrates agent submission, cancellation, timeout reconciliation, and SSE activity streaming into the execution workspace.
  • Adds production batch comparison controls with scoring preferences and MCP-backed results.
  • Adds activity, confirmation, mode-toggle, export, and catalog-notice UI components.
  • Extends batch-job timing data and supporting validation and feature-gating behavior.

Confidence Score: 4/5

The PR is not yet safe to merge because reconnecting an agent stream still appends an unchanged metrics snapshot as new activity.

The server replays logs before its latest metrics snapshot, while the client disables replay mode after matching the final log; the subsequent unchanged metrics event therefore bypasses deduplication and is appended again on each reconnect.

Files Needing Attention: src/components/features/execute/useAgentStream.ts

Important Files Changed

Filename Overview
src/components/features/execute/useAgentMode.ts Implements generation-aware agent submission, cancellation, startup timeout, deferred-stop, and orphan reconciliation; previously reported lifecycle races appear repaired.
src/components/features/execute/useAgentStream.ts Adds endpoint-specific SSE normalization and replay handling, but the previously reported reconnect path still duplicates the latest metrics snapshot after log-prefix replay completes.
src/components/features/execute/BatchProcessingPanel.tsx Connects completed batch jobs to the comparison view and MCP compare_results flow, resolving the prior production-wiring findings.
src/components/features/execute/BatchComparisonView.tsx Adds comparison eligibility, scoring controls, winner presentation, exclusions, sorting, and accessible disabled-state messaging.
src/components/features/execute/ExecutionWorkspace.tsx Integrates agent mode, activity streaming, confirmation-based mode switching, report export, and the new workspace controls.

Reviews (38): Last reviewed commit: "fix: stale generation's grace-timer clea..." | Re-trigger Greptile

Comment thread src/components/features/execute/useAgentMode.ts Outdated
Comment thread src/components/features/execute/useAgentStream.ts
Comment thread src/components/features/execute/BatchComparisonView.tsx

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8a805d64b3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/components/features/execute/useAgentStream.ts
Comment thread src/components/features/execute/useAgentStream.ts Outdated
Comment thread src/components/features/execute/useAgentStream.ts Outdated
Comment thread src/components/features/execute/useAgentMode.ts Outdated
Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
Comment thread src/components/features/execute/ExportReportMenu.tsx

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds a new “Agent” execution mode to the Execute workspace, including an activity log fed by SSE, controls + confirmation dialog for stopping the agent, and an expanded batch comparison UI (2–10 jobs) with scoring preference and winner highlighting.

Changes:

  • Introduces Agent mode UI building blocks (ModeToggle, AgentControls, AgentConfirmDialog) and integrates them into ExecutionWorkspace.
  • Adds agent activity logging components/hooks (ActivityLog, ActivityLogEntry, useAgentMode, useAgentStream) plus thorough unit/component tests.
  • Expands BatchComparisonView with scoring preference selection, compare-results rendering (MCP output), and winner/exclusion presentation.

Reviewed changes

Copilot reviewed 18 out of 18 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
src/components/features/input/CatalogUpdateNotice.tsx Adds a staleness notice with refresh CTA for recipe catalog updates.
src/components/features/execute/useAgentStream.ts Implements SSE hook with retry/backoff for agent activity streaming.
src/components/features/execute/useAgentStream.test.ts Unit tests for SSE parsing, lifecycle, and reconnection logic.
src/components/features/execute/useAgentMode.ts Local state hook for agent mode lifecycle, log buffering, and outcomes.
src/components/features/execute/useAgentMode.test.ts Unit tests for agent mode state transitions and timeout behavior.
src/components/features/execute/ModeToggle.tsx Segmented Manual/Agent mode toggle component.
src/components/features/execute/ModeToggle.test.tsx Component tests for the mode toggle’s behavior and ARIA roles.
src/components/features/execute/ExportReportMenu.tsx Adds feature-flagged report export dropdown (Markdown/PDF) with test IDs.
src/components/features/execute/ExecutionWorkspace.tsx Integrates agent-mode UI into Execute workspace and gates manual-only overlays.
src/components/features/execute/BatchComparisonView.tsx Extends comparison UI to 2–10 jobs, scoring preference, winner/exclusions, and MCP metrics table.
src/components/features/execute/BatchComparisonView.test.tsx Adds tests for new comparison behaviors (scoring, enable/disable, winner/exclusions, metrics).
src/components/features/execute/AgentControls.tsx Start/Stop buttons and agent status indicator component.
src/components/features/execute/AgentControls.test.tsx Component tests for agent controls states, outcomes, and accessibility.
src/components/features/execute/AgentConfirmDialog.tsx Confirmation dialog for switching away from agent mode while running.
src/components/features/execute/AgentConfirmDialog.test.tsx Component tests for dialog rendering, interactions, and keyboard/backdrop dismiss.
src/components/features/execute/ActivityLogEntry.tsx Renders individual activity entries with kind styling and truncation expand/collapse.
src/components/features/execute/ActivityLog.tsx Scrollable activity log with conditional auto-scroll and empty state.
src/components/features/execute/ActivityLog.test.tsx Component tests for rendering, accessibility, className, and auto-scroll behavior.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/components/features/execute/useAgentStream.ts
Comment thread src/components/features/execute/useAgentMode.ts
Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
@qodo-code-review

qodo-code-review Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Code Review by Qodo

🐞 Bugs (1) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Action required

1. ExportReportMenu has boolean props ✓ Resolved 📘 Rule violation ⌂ Architecture
Description
ExportReportMenuProps defines multiple boolean props (includeRecipeJson, includeLogSummary,
etc.) that customize behavior, violating the rule against boolean-prop-driven modes/variants. This
makes the component harder to extend without combinatorial prop logic.
Code

src/components/features/execute/ExportReportMenu.tsx[R25-28]

+  includeRecipeJson?: boolean;
+  /** Whether to include log summary in the report. */
+  includeLogSummary?: boolean;
+  /** Force-disable the menu (in addition to the zero-records check). */
Relevance

●●● Strong

Boolean-props anti-pattern has accepted precedent; likely refactor to enum/structured options prop.

PR-#117

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
PR Compliance ID 2436729 disallows using two or more boolean props to customize component behavior
or toggle rendering modes. ExportReportMenuProps introduces multiple boolean props controlling how
the menu behaves/what it includes.

Rule 2436729: Do not use boolean props to customize component behavior or toggle rendering modes
src/components/features/execute/ExportReportMenu.tsx[21-30]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`ExportReportMenu` uses multiple boolean props (`includeRecipeJson`, `includeLogSummary`, `disabled`) to control behavioral variants. The compliance rule requires avoiding boolean props for mode/variant switching when there are 2+ such booleans.

## Issue Context
These booleans encode multiple report "modes" and can lead to unclear/invalid combinations.

## Fix Focus Areas
- src/components/features/execute/ExportReportMenu.tsx[21-39]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Agent stream endpoint missing jobId parameter ✓ Resolved 🐞 Bug ≡ Correctness
Description
useAgentStream.ts opens an EventSource to a fixed /api/olive/agent/stream endpoint with no job
identifier and parses generic default-message JSON ({kind, text, stepRef}) plus a named 'done'
event, but the backend contract expects a job-scoped /agent/stream/:jobId URL and emits named
log, metrics, and done events. As a result, the hook cannot target a specific job and no
matching unscoped server route exists in this repo, so the agent activity stream cannot work
end-to-end.
Code

src/components/features/execute/useAgentStream.ts[17]

+const AGENT_STREAM_ENDPOINT = "/api/olive/agent/stream";
Relevance

●●● Strong

SSE endpoint/contract mismatch is a clear end-to-end break; team previously enforced job-scoped
stream usage.

PR-#97

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
In useAgentStream.ts, the constant AGENT_STREAM_ENDPOINT is passed directly into `new
EventSource(AGENT_STREAM_ENDPOINT)` and the hook’s options/signature provide no way to pass a jobId,
proving the client cannot scope the stream to a specific running job. Additionally, a repo-wide
search for the literal /api/olive/agent/stream only finds the hook and its test, indicating no
server route is registered at that exact unparameterized path, while existing repository behavior
expects a parameterized /agent/stream/:jobId URL and relies on named SSE events (log, metrics,
done) rather than only onmessage, demonstrating a protocol mismatch that prevents the UI from
populating the activity log.

src/components/features/execute/useAgentStream.ts[17-17]
src/components/features/execute/useAgentStream.ts[121-133]
src/components/features/execute/useAgentStream.ts[185-211]
src/server/routes/olive.ts[171-205]
src/server/routes/olive.ts[515-527]
src/components/features/execute/useOliveStream.ts[369-411]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`useAgentStream` currently connects to a fixed SSE endpoint (`/api/olive/agent/stream`) with no ability to scope the connection to a specific agent job, and it parses generic default `onmessage` JSON plus a `done` event. The backend/server contract in this repo expects a job-scoped stream URL (`/agent/stream/:jobId`) and emits named events (`log`, `metrics`, `done`), so the client cannot connect to the correct stream or translate events into activity updates.

## Issue Context
- The hook uses a literal endpoint constant and calls `new EventSource(...)` without any jobId parameterization, and `UseAgentStreamOptions` does not accept a job identifier.
- `ExecutionWorkspace.tsx` invokes the hook with `enabled: agentRunning` but does not thread a job ID into it.
- There is no server route registered at exactly `/api/olive/agent/stream` (unscoped) in this repository; existing UI execution code indicates the expected pattern is a parameterized URL and named-event handling.

## Fix Focus Areas
- src/components/features/execute/useAgentStream.ts[17-17]
- src/components/features/execute/useAgentStream.ts[121-133]
- src/components/features/execute/useAgentStream.ts[185-211]
- src/components/features/execute/ExecutionWorkspace.tsx[698-733]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Controls never orchestrate backend ✓ Resolved 🐞 Bug ≡ Correctness
Description
The Start and Stop buttons invoke hook functions that only mutate local React state; they neither
submit backend work nor call the agent cancellation route. The central Agent controls therefore
cannot start or cancel an actual execution.
Code

src/components/features/execute/ExecutionWorkspace.tsx[R791-794]

+            <AgentControls
+              agentRunning={agentRunning}
+              onStart={startAgent}
+              onStop={stopAgent}
Relevance

●●● Strong

PR intent says start/stop execution; prior reviews push real cancel/terminal wiring, not state-only
stubs.

PR-#97

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
ExecutionWorkspace passes the local hook callbacks directly to the buttons, and both hook
implementations only update state. The server exposes separate submission and cancellation APIs that
are never called by this path.

src/components/features/execute/ExecutionWorkspace.tsx[680-733]
src/components/features/execute/ExecutionWorkspace.tsx[791-794]
src/components/features/execute/useAgentMode.ts[111-167]
src/server/routes/olive.ts[445-491]
src/server/routes/olive.ts[569-595]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Agent controls currently change only local UI state and do not orchestrate backend work.

## Issue Context
Start must create/submit the agent execution and retain its job or session identifier; Stop must cancel that active backend execution before recording the local outcome.

## Fix Focus Areas
- src/components/features/execute/ExecutionWorkspace.tsx[791-794]
- src/components/features/execute/useAgentMode.ts[111-167]
- src/server/routes/olive.ts[445-491]
- src/server/routes/olive.ts[569-595]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View high (3)
4. CatalogUpdateNotice has boolean props ✓ Resolved 📘 Rule violation ⌂ Architecture
Description
CatalogUpdateNoticeProps introduces multiple boolean props (isStale, isRefreshing) that drive
rendering branches, violating the boolean-props rule. Consolidating into a single status/variant
prop avoids ambiguous combinations and simplifies usage.
Code

src/components/features/input/CatalogUpdateNotice.tsx[R5-8]

+  /** Whether the catalog is stale (upstream SHA differs from stored). */
+  isStale: boolean;
+  /** Whether a catalog refresh is currently in progress. */
+  isRefreshing: boolean;
Relevance

●●● Strong

Same boolean-props rule as prior accepted refactor; likely consolidate into single status/variant
prop.

PR-#117

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
PR Compliance ID 2436729 requires avoiding boolean-prop-driven modes/variants when a component has
two or more boolean props controlling behavior. CatalogUpdateNoticeProps adds both isStale and
isRefreshing, which directly control whether the component renders and which UI variant it shows.

Rule 2436729: Do not use boolean props to customize component behavior or toggle rendering modes
src/components/features/input/CatalogUpdateNotice.tsx[4-15]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
`CatalogUpdateNotice` uses multiple boolean props (`isStale`, `isRefreshing`) to drive rendering variants (hidden vs visible; spinner vs button). The compliance rule forbids 2+ boolean props for mode/variant toggling.

## Issue Context
The UI state can be represented as a single discriminated status (e.g., `status: 'upToDate' | 'stale' | 'refreshing'`) to avoid boolean combinations.

## Fix Focus Areas
- src/components/features/input/CatalogUpdateNotice.tsx[4-33]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


5. Comparison UI remains unreachable 🐞 Bug ≡ Correctness
Description
The enhanced BatchComparisonView has no production caller; JobHistoryModal still renders its own
legacy comparison and caps selection at six. Users cannot access the added scoring, winner,
excluded-job, or ten-job functionality.
Code

src/components/features/execute/BatchComparisonView.tsx[R90-96]

+export function BatchComparisonView({
+  records,
+  onClose,
+  compareResults,
+  onCompare,
+  completedJobCount,
+}: BatchComparisonViewProps) {
Relevance

●●● Strong

If UI isn’t mounted anywhere, reviewers usually request wiring or removal to avoid dead code.

PR-#58

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Repository-wide symbol search finds BatchComparisonView only in its declaration and tests. The
production workspace mounts JobHistoryModal, whose implementation retains a six-item selector and
separate card comparison.

src/components/features/execute/BatchComparisonView.tsx[90-100]
src/components/features/execute/ExecutionWorkspace.tsx[1304-1319]
src/components/features/execute/JobHistoryModal.tsx[141-145]
src/components/features/execute/JobHistoryModal.tsx[192-209]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The new comparison component is not integrated into the production history flow.

## Issue Context
Compose or replace the legacy comparison section in JobHistoryModal with BatchComparisonView and update selection limits to the supported 2–10 range.

## Fix Focus Areas
- src/components/features/execute/BatchComparisonView.tsx[90-96]
- src/components/features/execute/JobHistoryModal.tsx[141-145]
- src/components/features/execute/JobHistoryModal.tsx[192-209]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


6. Failed runs marked successful ✓ Resolved 🐞 Bug ≡ Correctness
Description
handleAgentStreamComplete always creates a successful outcome, even though the server's done
payload identifies completed, failed, and cancelled jobs. Any terminal failure or cancellation
received through the stream is consequently displayed as a completed Agent run.
Code

src/components/features/execute/ExecutionWorkspace.tsx[R721-724]

+    completeAgent({
+      status: "success",
+      totalSteps: agentEntries.length,
+      elapsedMs: agentStartedAt ? Date.now() - new Date(agentStartedAt).getTime() : 0,
Relevance

●●● Strong

Strong precedent fixing misreported cancel/fail outcomes from SSE; hardcoded success will be
corrected.

PR-#97

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new callback hard-codes success, while the server includes status and exitCode in every
done event. The existing Olive stream consumer already parses that payload and maps
cancellation/failure separately.

src/components/features/execute/ExecutionWorkspace.tsx[720-726]
src/components/features/execute/useAgentStream.ts[203-212]
src/server/routes/olive.ts[199-205]
src/server/routes/olive.ts[233-239]
src/components/features/execute/useOliveStream.ts[404-422]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Agent stream completion discards the server's terminal status and always reports success.

## Issue Context
Parse the `done` event payload, pass its status through the hook callback, and map completed, failed, and cancelled states to the corresponding AgentOutcome.

## Fix Focus Areas
- src/components/features/execute/useAgentStream.ts[203-212]
- src/components/features/execute/ExecutionWorkspace.tsx[720-726]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

7. Wrong jobs validate comparison ✓ Resolved 🐞 Bug ≡ Correctness
Description
The Compare button validates the optional count of all available completed jobs instead of the
selected records that will be compared. Unrelated jobs can enable a one-record comparison, and an
oversized selected set can pass when the supplied global count remains within 2–10.
Code

src/components/features/execute/BatchComparisonView.tsx[R127-129]

+  // Determine effective completed job count for button state
+  const effectiveCompletedCount = completedJobCount ?? records.filter((r) => r.status === "completed").length;
+  const canCompare = validateJobCount(effectiveCompletedCount);
Relevance

●●● Strong

Deterministic button-enable bug (validates wrong count); likely fixed to use selected records
length.

PR-#58

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The prop is documented as an available completed-job count, while records is the actual comparison
set. The shared validator explicitly requires the input job count itself to be between two and ten.

src/components/features/execute/BatchComparisonView.tsx[28-37]
src/components/features/execute/BatchComparisonView.tsx[127-129]
src/components/features/execute/BatchComparisonView.tsx[176-181]
src/lib/batchComparison.ts[13-21]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Comparison eligibility is calculated from a count that may not describe the selected comparison input.

## Issue Context
Derive eligibility from the selected completed records passed to onCompare, and ensure only those records are submitted.

## Fix Focus Areas
- src/components/features/execute/BatchComparisonView.tsx[127-129]
- src/components/features/execute/BatchComparisonView.tsx[176-181]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


8. Winner hides excluded jobs ✓ Resolved 🐞 Bug ≡ Correctness
Description
The excluded-job list is nested under winner === null, although CompareResultsOutput permits
exclusions alongside a non-null winner. A comparison that selects a winner after excluding
incomplete jobs hides those exclusions and implies every requested job participated.
Code

src/components/features/execute/BatchComparisonView.tsx[R279-280]

+          {/* No clear winner + excluded jobs */}
+          {compareResults.winner === null && (
Relevance

●●● Strong

UI branch contradicts result type (winner and exclusions independent); likely adjust rendering to
always show exclusions.

PR-#58

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The exclusion markup is inside the null-winner branch, while the result type defines winner and
excluded_jobs as independent fields with no invariant tying exclusions to a null winner.

src/components/features/execute/BatchComparisonView.tsx[279-314]
src/lib/types/agentTypes.ts[107-115]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Excluded jobs disappear whenever the comparison also has a winner.

## Issue Context
Keep the no-winner notice conditional, but render any nonempty excluded_jobs list independently for both winner states.

## Fix Focus Areas
- src/components/features/execute/BatchComparisonView.tsx[279-314]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


9. New export/catalog UI components unreachable from workspace ✓ Resolved 🐞 Bug ⚙ Maintainability
Description
ExportReportMenu and CatalogUpdateNotice are added as fully functional, tested components in this
PR, but ExecutionWorkspace.tsx (the file this PR modifies to wire in the other new agent-mode
components) never imports or renders either of them, leaving their functionality unreachable from
the panel this PR updates.
Code

src/components/features/execute/ExecutionWorkspace.tsx[R31-36]

+import { useAgentMode } from "./useAgentMode";
+import { useAgentStream } from "./useAgentStream";
+import { ModeToggle } from "./ModeToggle";
+import { AgentControls } from "./AgentControls";
+import { ActivityLog } from "./ActivityLog";
+import { AgentConfirmDialog } from "./AgentConfirmDialog";
Relevance

●●● Strong

Team tends to avoid adding fully-built but unused UI; likely wire into ExecutionWorkspace or remove.

PR-#162

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
ExecutionWorkspace.tsx's new import block adds useAgentMode, useAgentStream, ModeToggle,
AgentControls, ActivityLog, and AgentConfirmDialog, but omits ExportReportMenu and
CatalogUpdateNotice; a repo-wide search for ExportReportMenu/CatalogUpdateNotice usage finds only
their own definition files, with no import site anywhere in /src.

src/components/features/execute/ExecutionWorkspace.tsx[31-36]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
ExportReportMenu and CatalogUpdateNotice are fully implemented and tested but are not rendered anywhere in the application, making their functionality (report export UI, catalog staleness notice) unreachable by end users.

## Issue Context
ExecutionWorkspace.tsx is the file this PR updates to wire in the other new agent-mode components (ModeToggle, AgentControls, ActivityLog, AgentConfirmDialog), but it does not add ExportReportMenu or CatalogUpdateNotice, and no other file in the repository imports them either.

## Fix Focus Areas
- src/components/features/execute/ExecutionWorkspace.tsx[31-36] (add imports/usage for ExportReportMenu near the existing manual export menu)
- src/components/features/execute/ExportReportMenu.tsx[34-46] (confirm props expected by a real caller, e.g. job records source)
- src/components/features/input/CatalogUpdateNotice.tsx[26-33] (wire to useCatalogPin in the recipe catalog panel)

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View medium (1)
10. Retry budget resets forever ✓ Resolved 🐞 Bug ☼ Reliability
Description
Every successful EventSource open resets the retry counter, so repeated open-then-disconnect
cycles always restart at the first retry and never reach the three-retry error path. An unstable
stream can reconnect indefinitely while the UI remains running.
Code

src/components/features/execute/useAgentStream.ts[R215-217]

+      evtSource.onopen = () => {
+        if (!isMountedRef.current) return;
+        retryCountRef.current = 0;
Relevance

●● Moderate

Retry-budget semantics are designy; could be intentional to keep reconnecting. No close precedent
found.

PR-#97

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The counter is unconditionally reset on open, and the error handler uses that same counter to decide
whether retries are exhausted. Thus any reconnect that reaches open receives a fresh budget.

src/components/features/execute/useAgentStream.ts[214-218]
src/components/features/execute/useAgentStream.ts[240-257]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The retry limit is defeated whenever a connection opens before disconnecting.

## Issue Context
Use a session-wide retry budget, or reset the budget only after a defined stable period or meaningful stream progress.

## Fix Focus Areas
- src/components/features/execute/useAgentStream.ts[214-218]
- src/components/features/execute/useAgentStream.ts[240-257]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context
✅ Compliance rules (platform): 33 rules
Review mode: 🧠 Deep: This is a dense, broad feature change spanning agent execution state, EventSource reconnection, workspace integration, batch comparison logic, accessibility, and export flows, creating many independent opportunities for subtle defects.

Grey Divider

Tip of the day
💡 Did you know, you can type 'qodo, fix this' on a finding and the fix lands right on your PR

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread src/components/features/execute/ExportReportMenu.tsx Outdated
Comment thread src/components/features/input/CatalogUpdateNotice.tsx Outdated
Comment thread src/components/features/execute/ExecutionWorkspace.tsx Outdated
Comment thread src/components/features/execute/ExecutionWorkspace.tsx
Comment thread src/components/features/execute/useAgentStream.ts Outdated
Comment thread src/components/features/execute/BatchComparisonView.tsx
Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
Comment thread src/components/features/execute/useAgentStream.ts
Comment thread src/components/features/execute/ExecutionWorkspace.tsx
@tonythethompson
tonythethompson force-pushed the feat/execute-agent-mode-batch-comparison branch from 8a805d6 to 024f007 Compare August 13, 2026 14:05
@qodo-code-review

Copy link
Copy Markdown
Contributor

Qodo Fixer

✅ Committed (4) · ☑ Fixed (4)

Grey Divider

Commits pushed directly to this PR — no separate fix PR opened.

Process — 4 fixed
  • ☑ Fixed: ExportReportMenu has boolean props
  • ☑ Fixed: Agent stream endpoint missing jobId parameter
  • ☑ Fixed: CatalogUpdateNotice has boolean props
  • ☑ Fixed: Failed runs marked successful
  • ⏭ Skipped (2)

Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
Comment thread src/components/features/execute/ExportReportMenu.tsx
Comment thread src/components/features/execute/ExecutionWorkspace.tsx Outdated
Comment thread src/components/features/execute/ModeToggle.tsx Outdated
Comment thread src/components/features/execute/useAgentMode.ts
Comment thread src/components/features/execute/BatchProcessingPanel.tsx Outdated
Comment thread src/components/features/execute/ExecutionWorkspace.tsx Outdated
Comment thread src/components/features/execute/ExecutionWorkspace.tsx Outdated
Comment thread src/components/features/execute/BatchProcessingPanel.tsx Outdated
Comment thread src/components/features/execute/useAgentStream.ts Outdated
Comment thread src/components/features/execute/useAgentStream.ts Outdated
Comment thread src/components/features/execute/useAgentMode.ts Outdated
Comment thread src/components/features/execute/BatchComparisonView.tsx Outdated
Comment thread src/components/features/execute/useAgentMode.ts Outdated
Comment thread src/components/features/execute/ExecutionWorkspace.tsx Outdated
Comment thread src/components/features/execute/BatchProcessingPanel.tsx
Comment thread src/components/features/execute/useAgentMode.ts Outdated
@codefactor-io

codefactor-io Bot commented Aug 14, 2026

Copy link
Copy Markdown

CodeFactor found an issue: Error: Calling setState synchronously within an effect can trigger cascading renders

Effects are intended to synchronize state between React and external systems such as manually updating the DOM, state management libraries, or other platform APIs. In general, the body of an effect should do one or both of the following:

  • Update external systems with the latest state from React.
  • Subscribe for updates from some external system, calling setState in a callback function when external state changes.

Calling setState synchronously within an effect body causes cascading renders that can hurt performance, and is not recommended. (https://react.dev/learn/you-might-not-need-an-effect).

/app/src/components/features/execute/ExportReportMenu.tsx:87:29
85 |
86 | useEffect(() => {
> 87 | if (isDisabled && open) setOpen(false);
| ^^^^^^^ Avoid calling setState() directly within an effect
88 | }, [isDisabled, open]);
89 |
90 | // Hide entirely when the feature flag is disabled

It's currently on:
src\components\features\execute\ExportReportMenu.tsx:87
Commit b42339c

@codefactor-io

codefactor-io Bot commented Aug 14, 2026

Copy link
Copy Markdown

CodeFactor found an issue: Complex Method

It's currently on:
src\components\features\execute\BatchComparisonView.tsx:92-414
Commit b42339c

@codefactor-io

codefactor-io Bot commented Aug 14, 2026

Copy link
Copy Markdown

CodeFactor found an issue: Complex Method

It's currently on:
src\components\features\execute\useAgentStream.ts:303-343
Commit 6811dde

@codefactor-io

codefactor-io Bot commented Aug 14, 2026

Copy link
Copy Markdown

CodeFactor found an issue: Complex Method

It's currently on:
src\components\features\execute\BatchProcessingPanel.tsx:665-725
Commit b42339c

@codefactor-io

codefactor-io Bot commented Aug 14, 2026

Copy link
Copy Markdown

CodeFactor found an issue: Complex Method

It's currently on:
src\components\features\execute\useAgentMode.ts:231-404
Commit b42339c

@tonythethompson

Copy link
Copy Markdown
Owner Author

CodeFactor fixes

  • Merged duplicate `@/lib/jobHistoryStore` import in ExecutionWorkspace.tsx.
  • Wrapped `jobs` in its own `useMemo` in BatchProcessingPanel.tsx:770 so it no longer changes identity every render and destabilizes the downstream `comparisonRecords` memo.
  • Removed the synchronous `setState` inside an effect in ExportReportMenu.tsx:87 — derived `menuOpen = open && !isDisabled` at render time instead of syncing it back with an effect.

Left the two "Complex Method" findings (`BatchComparisonView.tsx`, `useAgentStream.ts`) as-is — those are cyclomatic-complexity heuristics on already-tested, working logic; splitting them is a refactor, not a fix, and out of scope here.

Typecheck clean (`tsc --noEmit`). Pushed in 590551e.

🤖 Addressed by Claude Code

Split handleCompare's fetch+parse into fetchCompareResults() and
startAgent's post-submit branching into resolveAgentSubmitResponse(),
same behavior, lower cyclomatic complexity per function.
@tonythethompson

Copy link
Copy Markdown
Owner Author

CodeFactor: Complex Method (round 2)

Extracted the two flagged functions without behavior changes:

  • BatchProcessingPanel.tsx `handleCompare`: pulled the fetch + response-parsing logic into a standalone `fetchCompareResults()` helper. `handleCompare` itself now just manages sequencing/abort/loading state.
  • useAgentMode.ts `startAgent`: pulled the post-submit stop/generation branching (the ~70-line block after the `/api/olive/jobs/submit` response) into a standalone `resolveAgentSubmitResponse()` helper, called and awaited from `startAgent`. Same logic, same throw-to-catch propagation, just out of the closure.

Typecheck clean, all 39 `useAgentMode`/`BatchProcessingPanel` tests pass. Pushed in c4b8216.

🤖 Addressed by Claude Code

BatchComparisonView: extract CompareControls, CompareResultsSection,
RecordsTable sub-components from the single render function.
useAgentStream: extract shouldDeliverStreamEntry() dedupe/replay check
from handleEntry. No behavior change.
@tonythethompson

Copy link
Copy Markdown
Owner Author

CodeFactor: Complex Method (round 3, remaining 2)

  • `BatchComparisonView.tsx`: split the single 320-line render function into `CompareControls`, `CompareResultsSection`, and `RecordsTable` sub-components. Main component now just composes them.
  • `useAgentStream.ts`: extracted the dedupe/replay-prefix decision out of `handleEntry` into `shouldDeliverStreamEntry()`.

No behavior change. Typecheck clean, 61 tests pass across `BatchComparisonView`/`useAgentStream`/`BatchProcessingPanel`. Pushed in 7232d63.

🤖 Addressed by Claude Code

The More-menu item downloaded ALL job history records unfiltered and
ignored the reportExport feature flag, diverging from ExportReportMenu
(completed-only, flag-gated). Removed the duplicate; ExportReportMenu
is the single source of truth.
onStop now appends an error log entry when stopAgent() returns false
instead of failing silently (agentRunning intentionally stays true so
Stop can be retried — confirmed design). handleAgentStreamComplete/Error
pass totalSteps: 0 so completeAgent()'s own stepCountRef (which only
counts stepRef entries) wins the Math.max instead of the inflated
agentEntries.length.
…etrics

useAgentMode: startAgent now generates an idempotencyKey and sends it
with /api/olive/jobs/submit. When armStopSubmitGrace's 30s grace timer
fires, it finalizes UI state exactly as before (unchanged timing), then
fire-and-forget re-POSTs with the same idempotencyKey in the background
— the server returns the existing job (reused: true) if one was created
before the abort landed, instead of spawning a duplicate, letting it be
found and cancelled instead of left orphaned.

useAgentStream: metrics events carry no id and are excluded from the
log-only prefix-replay check, so an unkeyed reconnect replayed the last
metrics snapshot as new every time. Added lastMetricsTextRef to dedupe
an exact repeat of the last delivered metrics text during a replay.
@tonythethompson

Copy link
Copy Markdown
Owner Author

Greptile: deferred-submission orphan risk + replayed metrics duplication

Both verified as real, still-open issues:

1. Deferred-stop grace timeout could leave an orphaned backend job. `armStopSubmitGrace`'s 30s grace timer only aborted the client-side request — if the server had already created the job before the abort landed (response lost), nothing reconciled it. Fixed by wiring up the server's existing `idempotencyKey` mechanism (`findJobByIdempotency` → `reused: true`): `startAgent` now generates and sends an idempotencyKey with the submit; when the grace timer fires, it finalizes UI state exactly as before (no timing change — all 16 existing timer-choreographed tests pass unmodified), then fire-and-forget re-POSTs with the same key in the background. If the server already created a job, the re-POST returns it instead of spawning a duplicate, so it can be found and cancelled.

2. Unkeyed replayed metrics bypassed reconnect dedup. Metrics events carry no SSE id and are intentionally excluded from the log-only prefix-replay check (comment: "Metrics between replayed logs must not desync the skip window") — but that also meant they had zero dedup, so every reconnect re-delivered the last metrics snapshot as a new activity entry. Added `lastMetricsTextRef` to skip an exact repeat of the last delivered metrics text during a replay.

Typecheck clean, all 160 execute-panel component tests pass (58 in useAgentMode/useAgentStream alone). Pushed in 43e8faa.

🤖 Addressed by Claude Code

Comment thread src/components/features/execute/useAgentMode.ts Outdated
…ession

clearStopSubmitGrace() operated on a single shared timer ref with no
ownership check. A stale, superseded generation's late-arriving submit
response (in resolveAgentSubmitResponse's thisGen !== runGenerationRef
branch) could clear a newer generation's currently-armed grace timer,
permanently stranding that session's stopAgent() waiter and blocking
the Agent-to-Manual confirmation dialog.

Tag the grace timer with its owning generation (stopSubmitGraceGenerationRef)
and only allow the stale-branch clear to proceed if it still owns the
armed timer.
@tonythethompson
tonythethompson merged commit a8ae813 into main Aug 14, 2026
13 checks passed
@tonythethompson
tonythethompson deleted the feat/execute-agent-mode-batch-comparison branch August 14, 2026 07:44
@linear-code

linear-code Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

OLI-118

tonythethompson added a commit that referenced this pull request Aug 14, 2026
…omponents (#304)

Break the ~1370-line ExecutionWorkspace orchestrator into feature-focused,
colocated files per the CodeRabbit plan, extending it to cover the agent-mode
section that landed with #285:

Hooks / pure utils:
- executionJobUtils: collectActivePassNames, resolveQueuedModelIdentifier,
  buildQueuedBatchJob, describeAppliedMcpPatches (+ unit tests)
- useDiagnosis: MCP diagnostic lifecycle, history, apply-fix mapping
- useOwrExport: OWR export overlay state + bundle download (fflate zip)
- useRecipeView: graph/json view state, More menu, export overlay actions

Presentational components:
- ExportRecipeOverlay, RecipePreviewCard, ExecutionLogPanel,
  ManualExecutionControls, AgentModeSection

ExecutionWorkspace.tsx remains a thin orchestrator wiring the hooks and
presentational sections. Behavior is unchanged: existing component tests pass
unmodified.
tonythethompson added a commit that referenced this pull request Aug 14, 2026
…omponents (#304) (#320)

* refactor: split ExecutionWorkspace into colocated hooks and feature components (#304)

Break the ~1370-line ExecutionWorkspace orchestrator into feature-focused,
colocated files per the CodeRabbit plan, extending it to cover the agent-mode
section that landed with #285:

Hooks / pure utils:
- executionJobUtils: collectActivePassNames, resolveQueuedModelIdentifier,
  buildQueuedBatchJob, describeAppliedMcpPatches (+ unit tests)
- useDiagnosis: MCP diagnostic lifecycle, history, apply-fix mapping
- useOwrExport: OWR export overlay state + bundle download (fflate zip)
- useRecipeView: graph/json view state, More menu, export overlay actions

Presentational components:
- ExportRecipeOverlay, RecipePreviewCard, ExecutionLogPanel,
  ManualExecutionControls, AgentModeSection

ExecutionWorkspace.tsx remains a thin orchestrator wiring the hooks and
presentational sections. Behavior is unchanged: existing component tests pass
unmodified.

* fix: align QNN ABI validation label

Co-authored-by: tonythethompson <32500316+tonythethompson@users.noreply.github.com>

* fix: resolve execution workspace review findings

* fix: bundle the desktop server so Linux package smoke can start

Packaged debs never shipped node_modules, but server.mjs was built with
external packages. Desktop builds now emit a self-contained server and
the smoke job no longer requires a packaged node_modules tree.

(cherry picked from commit 54c57ef)

* fix: support dynamic requires in desktop server bundle

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants