feat: add Execute agent mode, activity log, and batch comparison UI - #285
Conversation
Add core libraries for the v0.5 release: - src/lib/types: agent and finding type definitions - src/lib/featureFlags: safe boolean parsing with URL/localStorage/env precedence - src/lib/batchComparison: job count validation (2-10), finite number checks, MCP output structural validation - src/lib/multiLoraValidation: finite VRAM handling, adapter name/path/alpha validation, count limits - src/lib/activityLog: floor-based elapsed formatting, optional stepRef - src/lib/reportGenerator: correct default title, UTC timestamps, PDF export - src/lib/reviewReconciler: deterministic findings authoritative, AI findings targeting deterministic pass fields suppressed - src/lib/workspaceFingerprint: deterministic fingerprints excluding transient fields, stale async result handling via computation ID ref - src/lib/recipeCatalogPin: GitHub rate-limit handling, truncated tree rejection, content decoding hardening, stale detection with upstream SHA - src/lib/oliveRecipeBuilder: MultiLoRA gating behind PEFT + valid adapters, safe vramEstimateGb default - src/lib/actionExecutor: state mutations routed through commitUiStateUpdate - src/lib/hooks: useCatalogPin (refresh guard, resolve debounce, SHA propagation) and useWorkspaceFingerprint (memoization with stale result rejection) All libraries include unit and property-based tests. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
Sorry @tonythethompson, your pull request is larger than the review limit of 150000 diff characters
|
ⓘ Your Qodo trial ends soon. Ask your workspace admin to set up billing to keep reviews running after the trial. Manage billing |
|
Warning Review limit reached
Next review available in: 62 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (15)
📝 WalkthroughWalkthroughThe execution workspace now supports Manual and Agent modes, agent lifecycle and SSE activity streaming, bounded activity logs, confirmation dialogs, report export, and batch-result comparison. New catalog pinning, workspace fingerprinting, shared finding contracts, and property-based validation tests are also included. ChangesExecution features
Workspace infrastructure
Estimated code review effort: 5 (Critical) | ~120 minutes Mergeability Score: 🟡 Moderate · up to This PR adds agent execution and comparison/reporting flows, but the current head can duplicate activity entries after reconnects, show stale comparison winners after the job set changes, and leave comparison state stuck when response parsing stalls; dialog semantics and inconsistent report-export behavior also remain unresolved. These can produce misleading results or inaccessible controls, so merge should wait for fixes or explicit owner acceptance. Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 8✅ Passed checks (8 passed)
✨ Finishing Touches 💡 1⚔️ Resolve merge conflicts 💡
📝 Generate docstrings
🧪 Generate unit tests (beta)
✨ Simplify code
Warning Review ran into problems🔥 ProblemsLinked repositories: Public OSS repositories can only analyze public repositories installed in this organization. Analyzed Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
PR Summary by QodoAdd Agent execute mode with SSE activity log and enhanced batch comparison UI
AI Description
Diagram
High-Level Assessment
Files changed (18)
|
Greptile SummaryThis PR adds Execute agent mode, streamed activity, batch-run comparison, report export, and associated workspace controls.
Confidence Score: 4/5The PR is not yet safe to merge because reconnecting an agent stream still appends an unchanged metrics snapshot as new activity. The server replays logs before its latest metrics snapshot, while the client disables replay mode after matching the final log; the subsequent unchanged metrics event therefore bypasses deduplication and is appended again on each reconnect. Files Needing Attention: src/components/features/execute/useAgentStream.ts
|
| Filename | Overview |
|---|---|
| src/components/features/execute/useAgentMode.ts | Implements generation-aware agent submission, cancellation, startup timeout, deferred-stop, and orphan reconciliation; previously reported lifecycle races appear repaired. |
| src/components/features/execute/useAgentStream.ts | Adds endpoint-specific SSE normalization and replay handling, but the previously reported reconnect path still duplicates the latest metrics snapshot after log-prefix replay completes. |
| src/components/features/execute/BatchProcessingPanel.tsx | Connects completed batch jobs to the comparison view and MCP compare_results flow, resolving the prior production-wiring findings. |
| src/components/features/execute/BatchComparisonView.tsx | Adds comparison eligibility, scoring controls, winner presentation, exclusions, sorting, and accessible disabled-state messaging. |
| src/components/features/execute/ExecutionWorkspace.tsx | Integrates agent mode, activity streaming, confirmation-based mode switching, report export, and the new workspace controls. |
Reviews (38): Last reviewed commit: "fix: stale generation's grace-timer clea..." | Re-trigger Greptile
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8a805d64b3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
Pull request overview
Adds a new “Agent” execution mode to the Execute workspace, including an activity log fed by SSE, controls + confirmation dialog for stopping the agent, and an expanded batch comparison UI (2–10 jobs) with scoring preference and winner highlighting.
Changes:
- Introduces Agent mode UI building blocks (
ModeToggle,AgentControls,AgentConfirmDialog) and integrates them intoExecutionWorkspace. - Adds agent activity logging components/hooks (
ActivityLog,ActivityLogEntry,useAgentMode,useAgentStream) plus thorough unit/component tests. - Expands
BatchComparisonViewwith scoring preference selection, compare-results rendering (MCP output), and winner/exclusion presentation.
Reviewed changes
Copilot reviewed 18 out of 18 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| src/components/features/input/CatalogUpdateNotice.tsx | Adds a staleness notice with refresh CTA for recipe catalog updates. |
| src/components/features/execute/useAgentStream.ts | Implements SSE hook with retry/backoff for agent activity streaming. |
| src/components/features/execute/useAgentStream.test.ts | Unit tests for SSE parsing, lifecycle, and reconnection logic. |
| src/components/features/execute/useAgentMode.ts | Local state hook for agent mode lifecycle, log buffering, and outcomes. |
| src/components/features/execute/useAgentMode.test.ts | Unit tests for agent mode state transitions and timeout behavior. |
| src/components/features/execute/ModeToggle.tsx | Segmented Manual/Agent mode toggle component. |
| src/components/features/execute/ModeToggle.test.tsx | Component tests for the mode toggle’s behavior and ARIA roles. |
| src/components/features/execute/ExportReportMenu.tsx | Adds feature-flagged report export dropdown (Markdown/PDF) with test IDs. |
| src/components/features/execute/ExecutionWorkspace.tsx | Integrates agent-mode UI into Execute workspace and gates manual-only overlays. |
| src/components/features/execute/BatchComparisonView.tsx | Extends comparison UI to 2–10 jobs, scoring preference, winner/exclusions, and MCP metrics table. |
| src/components/features/execute/BatchComparisonView.test.tsx | Adds tests for new comparison behaviors (scoring, enable/disable, winner/exclusions, metrics). |
| src/components/features/execute/AgentControls.tsx | Start/Stop buttons and agent status indicator component. |
| src/components/features/execute/AgentControls.test.tsx | Component tests for agent controls states, outcomes, and accessibility. |
| src/components/features/execute/AgentConfirmDialog.tsx | Confirmation dialog for switching away from agent mode while running. |
| src/components/features/execute/AgentConfirmDialog.test.tsx | Component tests for dialog rendering, interactions, and keyboard/backdrop dismiss. |
| src/components/features/execute/ActivityLogEntry.tsx | Renders individual activity entries with kind styling and truncation expand/collapse. |
| src/components/features/execute/ActivityLog.tsx | Scrollable activity log with conditional auto-scroll and empty state. |
| src/components/features/execute/ActivityLog.test.tsx | Component tests for rendering, accessibility, className, and auto-scroll behavior. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Code Review by Qodo
1.
|
- Separate legacy adapter emission - Handle unknown VRAM conservatively
8a805d6 to
024f007
Compare
Qodo Fixer✅ Committed (4) · ☑ Fixed (4) Commits pushed directly to this PR — no separate fix PR opened. Process — 4 fixed
|
|
CodeFactor found an issue: Error: Calling setState synchronously within an effect can trigger cascading renders Effects are intended to synchronize state between React and external systems such as manually updating the DOM, state management libraries, or other platform APIs. In general, the body of an effect should do one or both of the following:
Calling setState synchronously within an effect body causes cascading renders that can hurt performance, and is not recommended. (https://react.dev/learn/you-might-not-need-an-effect). /app/src/components/features/execute/ExportReportMenu.tsx:87:29 It's currently on: |
|
CodeFactor found an issue: Complex Method It's currently on: |
|
CodeFactor found an issue: Complex Method It's currently on: |
|
CodeFactor found an issue: Complex Method It's currently on: |
|
CodeFactor found an issue: Complex Method It's currently on: |
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
CodeFactor fixes
Left the two "Complex Method" findings (`BatchComparisonView.tsx`, `useAgentStream.ts`) as-is — those are cyclomatic-complexity heuristics on already-tested, working logic; splitting them is a refactor, not a fix, and out of scope here. Typecheck clean (`tsc --noEmit`). Pushed in 590551e. 🤖 Addressed by Claude Code |
Split handleCompare's fetch+parse into fetchCompareResults() and startAgent's post-submit branching into resolveAgentSubmitResponse(), same behavior, lower cyclomatic complexity per function.
CodeFactor: Complex Method (round 2)Extracted the two flagged functions without behavior changes:
Typecheck clean, all 39 `useAgentMode`/`BatchProcessingPanel` tests pass. Pushed in c4b8216. 🤖 Addressed by Claude Code |
BatchComparisonView: extract CompareControls, CompareResultsSection, RecordsTable sub-components from the single render function. useAgentStream: extract shouldDeliverStreamEntry() dedupe/replay check from handleEntry. No behavior change.
CodeFactor: Complex Method (round 3, remaining 2)
No behavior change. Typecheck clean, 61 tests pass across `BatchComparisonView`/`useAgentStream`/`BatchProcessingPanel`. Pushed in 7232d63. 🤖 Addressed by Claude Code |
The More-menu item downloaded ALL job history records unfiltered and ignored the reportExport feature flag, diverging from ExportReportMenu (completed-only, flag-gated). Removed the duplicate; ExportReportMenu is the single source of truth.
onStop now appends an error log entry when stopAgent() returns false instead of failing silently (agentRunning intentionally stays true so Stop can be retried — confirmed design). handleAgentStreamComplete/Error pass totalSteps: 0 so completeAgent()'s own stepCountRef (which only counts stepRef entries) wins the Math.max instead of the inflated agentEntries.length.
…etrics useAgentMode: startAgent now generates an idempotencyKey and sends it with /api/olive/jobs/submit. When armStopSubmitGrace's 30s grace timer fires, it finalizes UI state exactly as before (unchanged timing), then fire-and-forget re-POSTs with the same idempotencyKey in the background — the server returns the existing job (reused: true) if one was created before the abort landed, instead of spawning a duplicate, letting it be found and cancelled instead of left orphaned. useAgentStream: metrics events carry no id and are excluded from the log-only prefix-replay check, so an unkeyed reconnect replayed the last metrics snapshot as new every time. Added lastMetricsTextRef to dedupe an exact repeat of the last delivered metrics text during a replay.
Greptile: deferred-submission orphan risk + replayed metrics duplicationBoth verified as real, still-open issues: 1. Deferred-stop grace timeout could leave an orphaned backend job. `armStopSubmitGrace`'s 30s grace timer only aborted the client-side request — if the server had already created the job before the abort landed (response lost), nothing reconciled it. Fixed by wiring up the server's existing `idempotencyKey` mechanism (`findJobByIdempotency` → `reused: true`): `startAgent` now generates and sends an idempotencyKey with the submit; when the grace timer fires, it finalizes UI state exactly as before (no timing change — all 16 existing timer-choreographed tests pass unmodified), then fire-and-forget re-POSTs with the same key in the background. If the server already created a job, the re-POST returns it instead of spawning a duplicate, so it can be found and cancelled. 2. Unkeyed replayed metrics bypassed reconnect dedup. Metrics events carry no SSE id and are intentionally excluded from the log-only prefix-replay check (comment: "Metrics between replayed logs must not desync the skip window") — but that also meant they had zero dedup, so every reconnect re-delivered the last metrics snapshot as a new activity entry. Added `lastMetricsTextRef` to skip an exact repeat of the last delivered metrics text during a replay. Typecheck clean, all 160 execute-panel component tests pass (58 in useAgentMode/useAgentStream alone). Pushed in 43e8faa. 🤖 Addressed by Claude Code |
…ession clearStopSubmitGrace() operated on a single shared timer ref with no ownership check. A stale, superseded generation's late-arriving submit response (in resolveAgentSubmitResponse's thisGen !== runGenerationRef branch) could clear a newer generation's currently-armed grace timer, permanently stranding that session's stopAgent() waiter and blocking the Agent-to-Manual confirmation dialog. Tag the grace timer with its owning generation (stopSubmitGraceGenerationRef) and only allow the stale-branch clear to proceed if it still owns the armed timer.
…omponents (#304) Break the ~1370-line ExecutionWorkspace orchestrator into feature-focused, colocated files per the CodeRabbit plan, extending it to cover the agent-mode section that landed with #285: Hooks / pure utils: - executionJobUtils: collectActivePassNames, resolveQueuedModelIdentifier, buildQueuedBatchJob, describeAppliedMcpPatches (+ unit tests) - useDiagnosis: MCP diagnostic lifecycle, history, apply-fix mapping - useOwrExport: OWR export overlay state + bundle download (fflate zip) - useRecipeView: graph/json view state, More menu, export overlay actions Presentational components: - ExportRecipeOverlay, RecipePreviewCard, ExecutionLogPanel, ManualExecutionControls, AgentModeSection ExecutionWorkspace.tsx remains a thin orchestrator wiring the hooks and presentational sections. Behavior is unchanged: existing component tests pass unmodified.
…omponents (#304) (#320) * refactor: split ExecutionWorkspace into colocated hooks and feature components (#304) Break the ~1370-line ExecutionWorkspace orchestrator into feature-focused, colocated files per the CodeRabbit plan, extending it to cover the agent-mode section that landed with #285: Hooks / pure utils: - executionJobUtils: collectActivePassNames, resolveQueuedModelIdentifier, buildQueuedBatchJob, describeAppliedMcpPatches (+ unit tests) - useDiagnosis: MCP diagnostic lifecycle, history, apply-fix mapping - useOwrExport: OWR export overlay state + bundle download (fflate zip) - useRecipeView: graph/json view state, More menu, export overlay actions Presentational components: - ExportRecipeOverlay, RecipePreviewCard, ExecutionLogPanel, ManualExecutionControls, AgentModeSection ExecutionWorkspace.tsx remains a thin orchestrator wiring the hooks and presentational sections. Behavior is unchanged: existing component tests pass unmodified. * fix: align QNN ABI validation label Co-authored-by: tonythethompson <32500316+tonythethompson@users.noreply.github.com> * fix: resolve execution workspace review findings * fix: bundle the desktop server so Linux package smoke can start Packaged debs never shipped node_modules, but server.mjs was built with external packages. Desktop builds now emit a self-contained server and the smoke job no longer requires a packaged node_modules tree. (cherry picked from commit 54c57ef) * fix: support dynamic requires in desktop server bundle --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Summary
ActivityLog+ActivityLogEntry: render agent stream entries with floor-based elapsed formattingAgentControls+useAgentMode: start/stop agent execution with mode toggle and confirmationAgentConfirmDialog: accessible confirmation dialog with test IDs andaria-describedbyuseAgentStream: EventSource hook with exponential backoff reconnection (1s, 2s, 4s) and 50ms grace period for done-event race handlingBatchComparisonView: compare 2-10 job records with sortable columns, duration deltas, scoring preference, winner highlighting, excluded jobs, and accessible disabled-control messagingExportReportMenu: report export with test IDsModeToggle: agent mode toggle with test IDCatalogUpdateNotice: catalog staleness noticeExecutionWorkspaceto integrate new componentsStacked on PR #282 (shared validation libs) which provides the underlying batch comparison, activity log, and agent type libraries.
Test plan
pnpm test:component- 239 component tests pass (32 test files), including BatchComparisonView (19), AgentConfirmDialog (10), ModeToggle (8), useAgentMode (15), useAgentStream (19)pnpm lint- 0 errorstsc --noEmit- 0 errorsGenerated with Devin
Note
Add Execute agent mode with activity log, batch comparison UI, and export report menu
ExecutionWorkspacewhere users can start/stop an autonomous agent session, stream real-time activity entries via SSE (useAgentStream), and switch back to manual mode with a confirmation dialog if the agent is running.useAgentModehook to manage agent session state (running flag, bounded activity log, outcome, timestamps, job ID) with a 10s start confirmation timeout.ActivityLog,ActivityLogEntry,AgentControls,AgentConfirmDialog, andModeTogglecomponents to support the agent mode UI.BatchComparisonViewandBatchProcessingPanelto support MCPcompare_resultscalls with a scoring preference selector, winner highlighting, and eligibility enforcement (2–10 completed jobs);BatchJobnow recordsstartedAtMs/finishedAtMstimestamps.ExportReportMenuto the workspace header (behind areportExportfeature flag) that downloads Markdown or triggers print-as-PDF for completed job records.Macroscope summarized 715cfe1.