Skip to content

feat(chat): show a Cowork run's real steps while it is still running - #1202

Merged
sakibsadmanshajib merged 5 commits into
mainfrom
feat/cowork-run-progress
Aug 26, 2026
Merged

sakibsadmanshajib merged 5 commits into
mainfrom
feat/cowork-run-progress

Conversation

@sakibsadmanshajib

Copy link
Copy Markdown
Owner

Follow-up to #1193, which made Cowork a mode of the composer and rendered a run as a conversation turn. Between submit and completion that turn showed nothing, which reads as a hang.

The gap was frontend only

The comment left in coworkMode.ts said the wire had no per-step feed. It has had one since #1073 merged:

  • agent-engine emits real per-step events, mapped from OpenHands ActionEvent, ObservationEvent, MessageEvent and AgentErrorEvent, with an unmapped class deliberately landing as a status event carrying the raw payload rather than being dropped.
  • control-plane stores them and serves GET /internal/agent-tasks/{id} and /events behind an after_seq cursor.
  • edge-api exposes GET /v1/agent/tasks/{id} and /events, FeatureCowork gated.
  • The chat proxy already routed get_task and list_task_events, with UUID and cursor validation.

vendor/open-webui/src/lib/hive/agentTasks.ts had listTasks, createTask and cancelTask and nothing else, so the run turn polled the whole task list and rendered no steps.

What this changes

agentTasks.ts: getTask and getTaskEvents, matching the proxy's contract. The cursor and limit are floored to plain non-negative integers, which is exactly what hive_agent_proxy.py accepts and the only shape that avoids a 400. decodeEvent keeps an event whose kind this build does not recognise, the way decodeTask already keeps an unrecognised status, because the backend goes out of its way not to drop an unmapped upstream class and dropping it at the last hop would undo that.

The follower: reads one task by id instead of listing every task the user owns and filtering it, which was the ponytail note in the code. Each poll then asks for events strictly after the highest seq the turn already carries. That seq is read off the stored lines, so a conversation reopened in another tab resumes where it left off instead of re-reading the run from zero. A full page means more is behind it and the loop asks again, bounded at five pages per read.

The rendering: events fold into muted lines on the same statusHistory field the chat path already uses for "Searching the web" and "Retrieved 3 sources". No new component, no branch in the transcript, and the treatment is the one the parity review found already matched the reference product. A tool call and its result join on tool_call_id into one line, which shimmers while the call is open.

Honesty properties, each with a test

  • Nothing renders that the backend did not send. No optimistic step, no synthesised progress. A status row that only repeats the task's own state contributes no line rather than a placeholder.
  • A payload the backend replaced with its truncation marker says so, and says the size. A preview sitting exactly on the 2000-rune cap is reported as shortened rather than shown as though complete: the cap leaves no marker behind, so the boundary is the only evidence there is.
  • A terminal run settles every line still open, so nothing shimmers under a turn that says the task finished. The text is not rewritten, because what is known is that the step stopped, not that it succeeded.
  • An event kind this build has never met still produces a line.
  • The events read is best effort and the status read is the authority: a failed event fetch keeps the lines already on screen and never erases real progress with a transport failure.

Not regressed, and pinned by tests

The three behaviours #1193's review fixed: the _chatId navigation guard, resuming every pending run rather than the oldest (selectPendingCoworkTurns), and the clean refusal when files are attached in Cowork mode.

Tests

scripts/test-owui-hive-frontend.sh: 175 passing, up from 147. Svelte compile pass green for all 13 Hive components; Chat.svelte compiled separately with the image build's pinned svelte 5.56.0 after TypeScript preprocessing.

Visual proof follows in a comment.

A run submitted from the composer showed nothing at all between submit and
completion, which reads as a hang. The reason recorded in the code was that the
wire had no per-step feed. It has had one since #1073: agent-engine emits
tool_call, tool_result, message, error, status and file events, control-plane
stores and serves them behind a cursor, edge-api exposes them under
GET /v1/agent/tasks/{id}/events, and the chat proxy already routed both that
and the by-id task read. The only missing piece was the frontend.

What changes:

  * agentTasks.ts gains getTask and getTaskEvents, matching the proxy's
    contract, including its refusal of anything that is not a plain
    non-negative integer cursor. Both new decoders keep a row whose kind or
    status this build does not recognise, the way decodeTask already did.
  * The run turn follows the detail endpoint instead of polling the whole task
    list and filtering it, which was the ponytail note left in the follower.
  * Events fold into muted transcript lines on the same statusHistory field the
    chat path already uses for "Searching the web", so per-step progress
    renders with no new component and no branch in the transcript. A tool call
    and its result join on tool_call_id into one line that shimmers while the
    call is open.
  * The cursor is used. Each poll asks only for events after the highest seq
    the turn already carries, read off the stored lines so a reopened
    conversation resumes where it left off rather than re-reading the run.

Honesty properties, all tested:

  * Nothing is rendered that the backend did not send. There is no optimistic
    step and no synthesised progress.
  * A payload the backend replaced with its truncation marker says so, with the
    size. A preview sitting exactly on the 2000-rune cap is reported as
    shortened rather than shown as complete.
  * A terminal run settles every line still open, so nothing shimmers under a
    turn that says the task finished.
  * An event kind this build has never met still produces a line, mirroring the
    syncer's deliberate refusal to drop an unmapped upstream class.

The three behaviours #1193's review fixed are unchanged and pinned by tests:
the _chatId navigation guard, resuming every pending run rather than the
oldest, and the clean refusal when files are attached in Cowork mode.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 28 minutes.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3934468c-7dff-43dd-b9a9-d94a5fa67374

📥 Commits

Reviewing files that changed from the base of the PR and between 7128e30 and 1508ab5.

📒 Files selected for processing (6)
  • docs/proof/cowork-run-progress-2026-08-26/capture-log.md
  • vendor/open-webui/src/lib/components/chat/Chat.svelte
  • vendor/open-webui/src/lib/hive/agentTasks.test.ts
  • vendor/open-webui/src/lib/hive/agentTasks.ts
  • vendor/open-webui/src/lib/hive/coworkMode.test.ts
  • vendor/open-webui/src/lib/hive/coworkMode.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Records the substrate honestly: the frontend is this branch built the way the
deploy builds it, and the agent sandbox plus the control-plane event syncer are
stood in for locally because the sandbox image is linux/amd64 Apptainer and
cannot run on this box. Every wire shape the stand-in serves is transcribed
from the shipped Go, with file references, and the browser's own request log is
included as the evidence that the follower reads one task by id and advances
the events cursor rather than re-reading the run.
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Visual proof

A run in flight, on this branch's own build. 1) t+28s: one muted line, shimmering, over the turn; before this change the same moment showed only "Working on it." 2) t+58s, expanded: seven lines, a tool call joined to its result, and the backend's truncation marker reported honestly. 3) t+65s: reloaded mid-run; the lines survive and the follower resumes from the stored cursor rather than re-reading the run. 4) t+93s: settled, nothing left shimmering. Substrate and its limits: docs/proof/cowork-run-progress-2026-08-26/capture-log.md

pr1202-20260826004754-31264-04-in-flight-t28s.png

pr1202-20260826004757-8137-06-in-flight-expanded-t58s.png

pr1202-20260826004801-14357-07-reloaded-mid-run-t65s.png

pr1202-20260826004804-32465-08-settled-t93s.png

…ered lines

An event that renders no line, which is what a status row repeating the task's
own state does and most rows are exactly that, left the cursor sitting behind
it. The next poll then asked for the same window again and kept re-reading that
tail for the rest of the run. Nothing was lost, since a later event still
arrived inside the same page, but the point of the cursor is that a poll costs
one small page, and this quietly gave that back. The page's own highest seq is
what was acknowledged, so that is what the cursor now carries; the stored lines
are still what a reopened conversation resumes from.
The first capture predated the cursor commit. Rendering is identical between
the two builds, but the committed request log is the evidence for how the
cursor behaves, so it is the log from the build that actually ships.
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Visual proof

Recaptured on the FINAL build of this branch, after the cursor commit (96c17d4); this supersedes the set above, which came from the build before it and renders identically. Same run, same local stack. 1) t+28s: one muted line, shimmering, over the turn. 2) t+58s, expanded: seven lines, a tool call joined to its result on tool_call_id, and the backend's truncation marker reported with its size. 3) t+70s: reloaded mid-run; every line survives, the run is still in flight, and the follower resumes at after_seq=10 rather than re-reading from zero. 4) t+99s: settled, nothing left shimmering, the run's own summary as the turn. Substrate and its limits, including why no live agent run is possible here: docs/proof/cowork-run-progress-2026-08-26/capture-log.md

pr1202-20260826010343-28620-04-in-flight-t28s.png

pr1202-20260826010346-6661-06-in-flight-expanded-t58s.png

pr1202-20260826010349-22552-07-reloaded-mid-run-t70s.png

pr1202-20260826010352-21002-08-settled-t99s.png

A step renders as one clamped line, so a marker appended to a 2000-rune
preview is the first thing the clamp throws away, and the line then reads as a
complete tool result. That is precisely the failure the marker exists to
prevent. The first capture of this surface showed it happening: a shortened
bash result rendered as "Used execute_bash:..." with no indication that
anything had been cut. The marker now sits next to the label, where the clamp
cannot reach it, and a message with no label carries it in front of the text.
@sakibsadmanshajib
sakibsadmanshajib merged commit a1ac2ef into main Aug 26, 2026
26 checks passed
@sakibsadmanshajib
sakibsadmanshajib deleted the feat/cowork-run-progress branch August 26, 2026 01:13
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Visual proof

Live capture against the deployed box (bf8edcb), real socket-arm sandbox, not the stub this PR originally captured against. Toggle/draft-preservation/launch/settle/reload-resume all hold. Per-step progress does NOT hold against a real run: 58s of real ls -la + file-write work produced zero new events (stuck at after_seq=6), and zero tool_call/tool_result/error events appeared at all. Full detail: PR #1205 (docs/proof/cowork-run-progress-live-2026-08-26/capture-log.md).

pr1202-20260826015136-31106-00-chat-loaded.png

pr1202-20260826015139-5257-03-cowork-mode-draft-check.png

pr1202-20260826015141-21765-05-submitted-t0.png

pr1202-20260826015144-6210-06-inflight-t35s.png

pr1202-20260826015146-23822-07-reloaded-t35s.png

pr1202-20260826015147-24167-09-final.png

sakibsadmanshajib added a commit that referenced this pull request Aug 26, 2026
…b) (#1205)

## Summary

- Verifies #1202's per-step progress rendering and #1193's composer mode
against a real deployed sandbox run, not the local `agent_stub.py`
#1202's own capture disclosed using (the three blockers that stub named
are gone: the demo box now carries #1193/#1202/#1203, the box is not
WSL2, and a live session could be minted).
- Toggle, draft preservation, real sandbox launch, settle behaviour, and
mid-run reload cursor resume all verified working.
- Per-step progress does **not** hold up against a real run: the
substantive 58-second work window produced zero new events, and the six
lines that did land are mostly dead text or noise. Zero
`tool_call`/`tool_result`/`error` events appeared despite real terminal
and file-editor tool use. Full detail in the capture log.

## Test plan

- [x] `node tools/lint-no-token-in-proof-captures.mjs` passes locally
against the new `docs/proof/cowork-run-progress-live-2026-08-26/`
directory
- [x] Screenshots posted as a follow-up comment on #1202 via
`scripts/post-pr-visual-proof.sh`
- [x] Findings verified independently against `public.agent_tasks` /
`public.agent_task_events` on the box's own Postgres, not only the UI

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Aug 26, 2026
## Diagnosis

Issue #1206, verified against the live stored rows for the run's own
task,
not inferred from code alone.

```
select task_id, seq, kind, length(payload::text), left(payload::text, 500)
from public.agent_task_events
where task_id = 'a98420c4-95a9-48cd-af03-3bea58e42241'
order by seq;
```

7 rows total, seq 1 through 7. Seq 1 (synthetic "running" status), 2
(SystemPromptEvent), 3 (the initiating user prompt, echoed), 4 and 5
(ConversationStateUpdateEvent), 6 (a bare `.git` workspace listing), 7
(synthetic "succeeded" status). Zero `tool_call`, zero `tool_result`,
zero
`file` event for `notes.md`, despite the run genuinely executing `ls
-la`
and writing `notes.md`.

This is a **mapping gap layered on an emission-sync gap**, and the sync
gap
is the one that actually produced the 58 second dead window:

- `EventSyncer.syncTask` (the per-pass sync for an active task) pulls
sandbox events and the workspace listing correctly. That produced seq 1
  through 6, all within the first few seconds of the run.
- `EventSyncer.finishVanished`, the path that runs once a task leaves
the
  active set (because the status `Poller` already recorded it terminal),
emitted **only** the bare terminal status event. It never called back
into
the sandbox events or workspace listing pull. Seq 7 is that bare status
row, and it is the *only* thing recorded between seq 6 (four seconds in)
  and task completion (82 seconds in). Whatever `ActionEvent` /
`ObservationEvent` pair the terminal command and the file write
produced,
wherever in that window OpenHands actually made them queryable, landed
in
exactly the pass `finishVanished` skips. That is the whole 58 second
gap:
  not five slow polls, one structurally incomplete final pass.
- Separately, real payload shapes at seq 2, 4, 5 show
`SystemPromptEvent`
  and `ConversationStateUpdateEvent` falling through `mapSandboxEvent`'s
unmapped-kind fallback. Their raw JSON has neither a `sandbox_kind` nor
a
  `status` top-level key (those are wrapper keys this repo's own code
introduces on the *other* branches), so the frontend's fallback renders
  "An update this version of Hive cannot read." for all three.

## Fix

- `EventSyncer.finishVanished` now shares the same
sandbox-events-plus-files
pull that an active pass uses (extracted to `pullTaskEvents`), so a
task's
tail is synced on its last possible pass instead of silently discarded.
  This is the root-cause fix, in the shared function both callers route
  through, not a rendering patch.
- `mapSandboxEvent` explicitly recognises and filters
`SystemPromptEvent`
and `ConversationStateUpdateEvent`: named, understood bookkeeping kinds,
excluded on purpose, not an unmapped kind silently vanishing. The
existing
"never drop an unknown kind silently" fallback is untouched for
genuinely
  unrecognised kinds.
- A user-role `MessageEvent` is now filtered too: the only way a user
  message enters this stream today is the task's own initiating prompt
(`InitialMessage` at launch); there is no running-task follow-up-message
  surface, so every one of these is that same prompt echoed back, not a
  real interjection.
- A dot-prefixed workspace entry (`.git`, ...) is filtered from the file
events: scaffolding, never agent output, and rendering it as a step was
  the other half of #1206's "reads as broken" complaint.

## Tests

Table-driven, over the real payload shapes captured from task `a98420c4`
(not invented shapes: PR #1202's guessed shape is exactly what let the
original defect ship). Plus a dedicated regression test,
`TestEventSyncerFinishVanishedSyncsTailEvents`, that reproduces the
failure
pattern directly: an empty first pass, then real `ActionEvent` /
`ObservationEvent` / file activity that only becomes visible on the pass
where the task has already left the active set, asserting all three
still
land instead of being dropped.

```
go build ./apps/control-plane/... ./apps/agent-engine/...
go vet ./apps/control-plane/... ./apps/agent-engine/...
go test ./apps/control-plane/internal/agenttask/... -count=1
```

All clean.

## Verification limits

Full live re-verification (launch a real Cowork run against the deployed
box and watch tool steps render during the run) requires this fix to
actually be running there, which only happens on merge to `main`
(`deploy-demo-box.yml`). This PR does not merge itself, so that step is
still pending. What's checked here instead: the fix is proven directly
against the real payload shapes that produced the defect, and against a
regression test that reproduces the exact tail-loss timing, not a
fixture
that merely resembles it.

Fixes #1206.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant