feat: Bot Screen — per-bot Xfce desktop streamed into Hermes Desktop, take over and hand back (related #92524) - #108914
Conversation
૮ >ﻌ< ა ci reviewran on 92e2e7c — Merge remote-tracking branch 'origin/main' into hermes/herme
|
| Package | Before | After |
|---|---|---|
| ➕ @novnc/novnc | — | 1.7.0 |
How to fix:
Add the ci-reviewed label after verifying the version changes are expected.
debug info
CI timings
CI timings · View report · View job
Wall time 4m49s vs 5m54s (-18.4%). 4 job(s) slower, 12 faster, 1 unchanged.
- Python tests / Run tests: -86.0s
- Python lints / Windows footguns (blocking): +48.0s
- OS-specific tests / Windows-only tests: -41.0s
- Docs Site / docs-site-checks: +20.0s
- OS-specific tests / macOS-only tests: -15.0s
d9525b3 to
3c70635
Compare
Julientalbot
left a comment
There was a problem hiding this comment.
Reviewing this from cognitive ergonomics and human–agent interaction, with AI-assisted source inspection and controlled regression probes, on cc0814d831b61f40183228da2fa2517487e018e3.
The useful advance here is continuity: a person can intervene in the agent's actual environment, then let it continue with the resulting browser session. Watch by default, named Take over / Hand back actions, visible ownership and server-side viewer-input filtering are valuable foundations.
Recommendation: address the control guarantees before approval. When someone takes over to enter credentials or complete 2FA, the product must prevent the agent from both acting on and observing the reserved surface. The attached comments identify five reproduced cases: browser dispatch under human control; a capture result crossing takeover and hand-back; abnormal disconnection releasing control; a stale preview still labelled Live; and control state applied to the wrong host.
The interaction contract I would look for is:
- Takeover: request → confirmed restriction of relevant agent access, with in-flight operations handled → human input enabled. Enforce this outside the model.
- Connection loss: keep agent access blocked and provide a recoverable interrupted state. A deadline may trigger recovery or escalation; elapsed time must not grant access. A watchdog is a possible implementation choice, not the requirement itself. Closing noVNC or rebooting does not establish that secrets have been erased.
- Hand-back: explicit transfer of control, followed by checking the resulting state. It does not certify successful authentication or completed work.
- Status: distinguish connection health, confirmed controller and task activity. “Agent can act” does not establish “agent has resumed.”
This builds on my published criteria for exclusive control and real delegation and keeping work state outside the user's working memory. The FAA's positive exchange of controls is a useful coordination analogy: an explicit, acknowledged transfer. Confidentiality is an additional requirement here.
Two design follow-ups: retain the handoff reason while the person acts—it is currently in a hover tooltip and cleared on acquisition—and make recovery preserve the access restriction across restarts. In particular, the lease reader defaults to agent ownership on missing/invalid state; legitimate initialization should be distinguished from loss of authority state during an active session. That last point is source inspection, not a sixth executed probe.
The documented shared OS account is not a security boundary. Fixing browser coordination alone cannot establish confidentiality against other agent-accessible routes such as terminal, files or child processes. The protection claim should match the enforced boundary. This does not make hosted-browser/dashboard support or per-bot OS users requirements of this PR; those are explicitly outside its scope.
Evidence and limits: the five probes are included under the inline comments. They exercise real handlers/components with doubles at stated execution or transport boundaries. No live Linux Xfce/noVNC trial or user study was conducted. The existing targeted Python suite had 12 passes and two Linux-only skips; the existing WebSocket URL tests had two passes. Installed local dependencies were reused. Predicted confusion and recovery effort remain hypotheses for live trials with the intended users.
| # Headed Chromium opens on this profile's Bot Desktop when one is running (human can take it over). | ||
| from tools.bot_desktop.runtime import desktop_env as _bot_desktop_env | ||
| return _bot_desktop_env(env) |
There was a problem hiding this comment.
[P1] Apply human control to the shared browser's actions and observations.
A person may take over to enter credentials while the bot still uses its ordinary browser tools. This change binds those tools to the same desktop and persistent browser profile, but their execution path does not check the human lease applied to computer_use. “You are in control” therefore does not establish the expected exclusion.
A probe through the actual browser_click handler, real environment construction and a persisted human lease dispatched agent-browser --session review --headed --engine chrome --json click @e1 and returned success. Only the final process execution and setup boundaries were stubbed; no real browser click was performed.
Please enforce the lease at the execution boundary for actions and reads of this shared local browser, and account for operations already in progress before confirming takeover. Keep unrelated remote browser sessions appropriately scoped. Acceptance: the shared browser cannot change or expose the human intervention through these agent tools while the human holds control.
Controlled regression probe — expected to fail on this head
Save the following review-only test as tests/tools/test_review_108914_browser_control.py in a disposable checkout. The test file is not part of this PR. Run scripts/run_tests.sh tests/tools/test_review_108914_browser_control.py -j 1 --file-retries 0.
import json
from tools.bot_desktop import lease, runtime
def test_browser_click_is_fenced_while_human_controls_shared_browser(monkeypatch):
from tools import browser_tool as browser
from tools import browser_tool_session as session
monkeypatch.setattr(runtime, "published_env", lambda: {"DISPLAY": ":37"})
monkeypatch.delenv("AGENT_BROWSER_PROFILE", raising=False)
assert browser._build_browser_env()["AGENT_BROWSER_PROFILE"].endswith("bot-desktop/browser-profile")
monkeypatch.setattr(browser, "_is_camofox_mode", lambda: False)
monkeypatch.setattr(browser, "_blocked_private_page_action", lambda *a: None)
monkeypatch.setattr(session, "_browser_command_preflight", lambda: {"browser_cmd": "agent-browser"})
monkeypatch.setattr(session, "_get_session_info", lambda *a: {"session_name": "review"})
monkeypatch.setattr(session._cloud, "_get_browser_engine", lambda: "chrome")
monkeypatch.setattr(session._cloud, "_is_headed_mode", lambda: True)
commands = []
def spawn(*args):
commands.append(args[2])
return {"success": True, "data": {}}
monkeypatch.setattr(session, "_spawn_and_collect", spawn)
lease.acquire("human-viewer")
result = json.loads(browser.browser_click("e1", task_id="review"))
assert commands == [], f"Human holds the lease, but the browser command was dispatched: {commands}; result={result}"There was a problem hiding this comment.
Fixed in 7a43ce8b57fe. _run_browser_command now brackets every local agent-browser command (no CDP url) with the lease while this profile's screen is running: refused up front with code: human_has_control, and a result whose run crossed a lease epoch change is discarded. Cloud/user-CDP sessions are untouched. Live: with a human holding the lease from another process, browser_click and browser_snapshot both returned human_has_control; navigate worked again after hand-back. Test: tests/tools/test_bot_desktop_browser_fence.py (your probe, red on the previous head).
| except _bd_lease.HumanHasControl as e: | ||
| return _refused(e) | ||
| result = _dispatch(backend, action, args) | ||
| if _bd_lease.get().epoch != admitted.epoch and _bd_lease.human_holds(): |
There was a problem hiding this comment.
[P1] Discard results across a completed takeover / hand-back cycle.
A capture admitted before takeover can produce a result during the human intervention and return after the human has already handed back. This condition only discards the result if the human still holds control, so a completed acquire → release cycle escapes the fence.
A controlled _dispatch seam performed acquire → synthetic capture result → release. The actual handle_computer_use returned the HUMAN_PRIVATE_FRAME marker. This proves the result-fencing gap without taking a real screenshot.
Please invalidate results crossing a relevant ownership epoch regardless of the final holder, with the admission epoch sampled at the execution boundary. Returning control authorizes subsequent observation; it should not authorize delivery of captures from the reserved human period.
Controlled regression probe — expected to fail on this head
Save the following review-only test as tests/tools/test_review_108914_capture_epoch.py in a disposable checkout. The test file is not part of this PR. Run scripts/run_tests.sh tests/tools/test_review_108914_capture_epoch.py -j 1 --file-retries 0.
import json
from tools.bot_desktop import lease, runtime
def test_capture_crossing_takeover_and_handback_discards_human_frame(monkeypatch):
from tools.computer_use import tool
monkeypatch.setattr(tool, "_get_backend", lambda session_id="": object())
def capture_during_handoff(backend, action, args):
lease.acquire("human-viewer")
captured = json.dumps({"ok": True, "action": action, "png_b64": "HUMAN_PRIVATE_FRAME"})
lease.release("human-viewer")
return captured
monkeypatch.setattr(tool, "_dispatch", capture_during_handoff)
result = tool.handle_computer_use({"action": "capture"})
assert "HUMAN_PRIVATE_FRAME" not in result, resultThere was a problem hiding this comment.
Fixed in 7a43ce8b57fe: epoch-only comparison, human_holds() dropped. test_takeover_during_an_admitted_action_discards_its_result now runs the full acquire+release cycle inside the dispatch and was red before the change.
| # Closing the viewer window hands control back; a stale holder never pins the agent out. | ||
| if _lease.viewer_may_send_input(viewer_id, profile_key=profile_home): | ||
| _lease.release(viewer_id, profile_key=profile_home) |
There was a problem hiding this comment.
[P1] Do not treat an abnormal disconnect as intentional hand-back.
The person may lose their connection halfway through entering credentials. This cleanup releases their lease regardless of the reason the stream ended. In the controlled endpoint probe below, a websocket.disconnect with code 1006 moved ownership to the agent and made wait_for_human return ok: true, without a Hand back action. Ticket consumption and the persisted lease were real; transport and the admission check were test doubles.
The tool does not literally claim the login succeeded. The problem is that connection loss grants renewed agent access to an unfinished human intervention. Closing the pane to hand back is documented, and avoiding a stranded lease is reasonable, but involuntary loss needs a distinct outcome.
Please preserve agent exclusion on connection loss and provide explicit recovery or release. A timeout can escalate recovery, but should not grant access. Acceptance: dropping the connection during a test login does not authorize new agent reads/actions; a reauthenticated person can recover and deliberately return control.
Controlled regression probe — expected to fail on this head
Save the following review-only test as tests/hermes_cli/test_review_108914_disconnect.py in a disposable checkout. The test file is not part of this PR. Run scripts/run_tests.sh tests/hermes_cli/test_review_108914_disconnect.py -j 1 --file-retries 0.
"""Review probe: connection loss must not imply an intentional hand-back.
Real endpoint, ticket consumption and persisted lease; only auth / transport are
test doubles. This is not a live noVNC or Linux desktop test.
"""
import asyncio
import json
from hermes_constants import get_hermes_home
from hermes_cli.dashboard_auth import ws_tickets
from hermes_cli.web_routers import display
from tools.bot_desktop import lease
from tools.computer_use.handoff import handle_handoff
def test_abnormal_viewer_disconnect_does_not_authorize_agent_to_resume(monkeypatch):
profile_home = get_hermes_home()
socket_path = profile_home / "bot-desktop" / "rfb.sock"
socket_path.parent.mkdir(parents=True, exist_ok=True)
socket_path.touch()
ws_tickets._reset_for_tests()
ticket = ws_tickets.mint_ticket(
user_id="display:review-human", provider="bot-desktop",
extra={"hermes_home": str(profile_home), "viewer_id": "review-human"},
)
lease.request_handoff("Finish login and 2FA")
lease.acquire("review-human")
class Viewer:
query_params = {"display_ticket": ticket}
async def accept(self):
pass
async def receive(self):
return {"type": "websocket.disconnect", "code": 1006}
async def close(self, **kwargs):
pass
class DesktopReader:
async def read(self, count):
await asyncio.Event().wait()
class DesktopWriter:
def close(self):
pass
async def connect(path):
assert path == str(socket_path)
return DesktopReader(), DesktopWriter()
monkeypatch.setattr(display, "_ws_request_is_allowed", lambda ws: True)
monkeypatch.setattr(display.asyncio, "open_unix_connection", connect)
asyncio.run(display.display_ws(Viewer()))
result = json.loads(handle_handoff("wait_for_human", {"seconds": 1}))
assert not result["ok"], f"No Hand back action occurred; abnormal disconnect produced: {result}"There was a problem hiding this comment.
Fixed in 7a43ce8b57fe. The bridge records the close code from websocket.disconnect; only 1000/1001 (viewer closed the window) release the lease. Anything else keeps the human's exclusion; the Desktop reconnects into the same lease or the human presses Hand back. Live against hermes serve: transport.abort() (1006) → holder stays human; close(1000) → agent. Test: tests/hermes_cli/test_display_ws_drop_keeps_lease.py (parametrised 1006/1000; the bridge body was lifted into _bridge(ws, info) so it is testable).
| .catch(() => { | ||
| /* transient: the next tick retries; the last good frame stays up */ | ||
| }) |
There was a problem hiding this comment.
[P2] Mark a retained frame as stale when refreshes fail.
A returning user may rely on “Live · bot in control” when deciding whether to keep watching. This catch retains the last image while the caption still derives from cached runtime/lease state. There is no failed-refresh or frame-age state to qualify “Live.”
The real hero component and portal-state hook were rendered with one successful thumbnail followed by rejected requests. After 60 seconds of simulated time, the button still said Screen: Live · bot in control. This was a component probe, not a real network outage; 60 seconds demonstrates persistence, not a proposed universal freshness threshold.
Please retain useful visual context but show when it was last received/captured, mark failed or stale updates, and reconcile status on recovery. Do not infer that the agent stopped merely because the viewing connection failed. Acceptance: repeated refresh failures cannot leave an old frame presented as unqualified live evidence.
Controlled regression probe — expected to fail on this head
Save the following review-only test as apps/desktop/src/plugins/hermes-bots/screen-review-108914-freshness.test.tsx in a disposable checkout. The test file is not part of this PR. Run cd apps/desktop && ../../node_modules/.bin/vitest run src/plugins/hermes-bots/screen-review-108914-freshness.test.tsx.
import { act, render } from '@testing-library/react'
import { expect, it, vi } from 'vitest'
vi.mock('@hermes/plugin-sdk', async () => {
const { useStore } = await import('@nanostores/react')
const { onGatewayEvent } = await import('../../contrib/events')
return { Codicon: () => null, useValue: useStore, host: { onEvent: onGatewayEvent } }
})
vi.mock('./data', async () => {
const { atom } = await import('nanostores')
return { $lastRoster: atom([]), botSelectionKey: (bot: { connectionId: string; name: string }) => `${bot.connectionId}:${bot.name}` }
})
vi.mock('./i18n', () => ({ useBots: () => ({ screen: {
portalTitle: 'Screen', portalWatching: 'Live · bot in control',
heroOpenLive: 'Open live', heroConnecting: 'Connecting',
} }) }))
vi.mock('./routing', () => ({ resolveBotConnectionRoute: () => ({ route: null }) }))
vi.mock('./screen-connection', () => ({ VIEWER_ID: 'review', displayRequest: vi.fn() }))
vi.mock('./screen-open', () => ({ openBotScreen: vi.fn() }))
import { ScreenHero } from './screen-hero'
import { displayRequest, type DisplayStatus } from './screen-connection'
import { $screenState, setScreenStatus } from './screen-state'
import type { RosterRow } from './types'
it('does not present a retained frame as live after repeated transport failures', async () => {
vi.useFakeTimers()
vi.spyOn(document, 'hidden', 'get').mockReturnValue(false)
const bot = { name: 'default', connectionId: 'host-a', sourceScoped: true } as RosterRow
$screenState.set({})
setScreenStatus(bot, {
profile: 'default', profile_key: '/home/hermes/.hermes', supported: true, installed: true, running: true,
lease: { holder: 'agent', viewer_id: null, pending_handoff: null, since: 1, reason: '' },
} as DisplayStatus)
const request = vi.mocked(displayRequest)
request.mockReset()
request.mockResolvedValueOnce({ data_url: 'data:image/jpeg;base64,LAST_GOOD_FRAME' })
.mockRejectedValue(new Error('Gateway unavailable'))
const view = render(<ScreenHero bot={bot} />)
try {
await act(async () => {})
expect(view.container.querySelector('img')?.getAttribute('src')).toContain('LAST_GOOD_FRAME')
expect(view.getByRole('button').getAttribute('aria-label')).toContain('Live')
// Fifteen consecutive failed refreshes: initial success cannot justify current "Live".
await act(async () => { await vi.advanceTimersByTimeAsync(60_000) })
expect(request.mock.calls.length).toBeGreaterThan(2)
expect(view.getByRole('button').getAttribute('aria-label')).not.toContain('Live')
} finally {
view.unmount()
vi.restoreAllMocks()
vi.useRealTimers()
}
})There was a problem hiding this comment.
Fixed in d947fc8: three consecutive failed refreshes dim the retained frame (grayscale, 40%) and caption it "Last seen — screen unreachable"; a success resets. Localised in all four locales.
| if (payload?.lease && payload.profile_key && payload.profile_key === profileKey) { | ||
| setScreenLease(bot, payload.lease) |
There was a problem hiding this comment.
[P2] Match the connection identity as well as the profile path.
Two hosts can both use /home/hermes/.hermes. The shared gateway event stream carries connectionId, but this listener matches only the path. A takeover event from host B can therefore replace host A's cached agent lease with B's human lease. Since VIEWER_ID is per desktop window, A can appear controlled by this window although its backend still permits its agent to act.
The probe below uses the actual event bus, useScreenPortalState and screen store: seed A as agent-controlled, emit B's human lease at the same path, and A's holder becomes human. It establishes false displayed ownership, not mixed video, transferred cookies or input delivered to another VM.
Please use the same connection/profile routing identity as display requests across screen event consumers, including the pane listener and installation events. Acceptance: an event from B cannot change A's control or installation state when their remote paths happen to match.
Controlled regression probe — expected to fail on this head
Save the following review-only test as apps/desktop/src/plugins/hermes-bots/screen-review-108914.test.tsx in a disposable checkout. The test file is not part of this PR. Run cd apps/desktop && ../../node_modules/.bin/vitest run src/plugins/hermes-bots/screen-review-108914.test.tsx.
import { act, renderHook } from '@testing-library/react'
import { expect, it, vi } from 'vitest'
vi.mock('@hermes/plugin-sdk', async () => {
const { useStore } = await import('@nanostores/react')
const { onGatewayEvent } = await import('../../contrib/events')
return { Codicon: () => null, useValue: useStore, host: { onEvent: onGatewayEvent } }
})
vi.mock('./data', async () => {
const { atom } = await import('nanostores')
return { $lastRoster: atom([]), botSelectionKey: (bot: { connectionId: string; name: string }) => `${bot.connectionId}:${bot.name}` }
})
vi.mock('./i18n', () => ({ useBots: () => ({}) }))
vi.mock('./routing', () => ({ resolveBotConnectionRoute: () => ({ route: null }) }))
vi.mock('./screen-connection', () => ({ VIEWER_ID: 'same-desktop-window', displayRequest: vi.fn() }))
vi.mock('./screen-open', () => ({ openBotScreen: vi.fn() }))
import { emitGatewayEvent } from '../../contrib/events'
import { useScreenPortalState } from './screen-portal'
import { $screenState, setScreenStatus } from './screen-state'
import type { DisplayStatus } from './screen-connection'
import type { RosterRow } from './types'
it('keeps host A control state unchanged when host B has the same home path', () => {
const botA = { name: 'default', connectionId: 'host-a', sourceScoped: true, connectionKind: 'remote' } as RosterRow
const status = {
profile: 'default', profile_key: '/home/hermes/.hermes', supported: true, installed: true, running: true,
lease: { holder: 'agent', viewer_id: null, pending_handoff: null, since: 1, reason: '' }
} as DisplayStatus
$screenState.set({})
setScreenStatus(botA, status)
const { result, unmount } = renderHook(() => useScreenPortalState(botA))
expect(result.current.lease?.holder).toBe('agent')
act(() => emitGatewayEvent({
type: 'display.lease', connectionId: 'host-b', profile: 'default',
payload: { profile_key: status.profile_key, lease: { holder: 'human', viewer_id: 'same-desktop-window', pending_handoff: null, since: 2, reason: '' } }
}))
try {
expect(result.current.lease?.holder).toBe('agent')
} finally {
unmount()
}
})There was a problem hiding this comment.
Fixed in d947fc8: one predicate isEventForBotScreen(bot, event, profileKey) in screen-connection.ts requires event.connectionId to equal the bot's route connection (untagged events only match a local route) AND the profile key; used by the pane lease listener and both install listeners. Test: screen-connection.test.ts.
cc0814d to
d947fc8
Compare
|
Thanks @Julientalbot — every finding reproduced. All five plus both design notes are addressed in 7a43ce8b57fe and d947fc8 (head d947fc8); inline replies carry the per-item evidence. The two design notes: the lease reader now fails CLOSED on a present-but-unparsable file (only a missing file is a fresh agent-held profile; the next successful write repairs it), and a takeover after Live re-verification on a real Xvnc screen against |
carlotestor
left a comment
There was a problem hiding this comment.
Tested this on a headless host (no X packages, so the noVNC/takeover UI is untested — this is the Python half only).
What passed: the PR's own suite is 19/19 green. I also wrote 13 independent adversarial tests against RfbClientFilter, since server-side input filtering is the load-bearing security claim here, and it holds up well:
- KeyEvent / PointerEvent / ClientCutText / QEMU Extended KeyEvent all dropped without the lease; FramebufferUpdateRequest still passes
- byte-by-byte fragmentation gives byte-identical output to bulk feeding, and 50 randomized chunkings of a 120-message stream leaked zero input bytes — nothing sneaks through a split message
- variable-length framing (SetEncodings, Fence, extended clipboard with negative length) stays correctly framed, so an input message can't hide behind a mis-parsed one
- lease flip applies to the very next message; ClientInit shared-flag is forced to 1 even when the handshake is split across chunks
Lease under real concurrency is clean too: 6 processes × 200 acquire/release cycles on one file ended holder=agent, epoch 1970, no torn or corrupt state, no leftover .tmp files; a stale viewer's release doesn't yank control from the current holder.
One finding: unbounded memory growth in the RFB parser, reachable by a viewer with no lease.
_message_length() trusts the client-declared ClientCutText length, and feed() accumulates into self._buf until that many bytes arrive. Nothing caps _buf, and display.py:124 feeds every WebSocket frame in with no limit of its own. So a read-only watcher — someone who only ever had view access and never held the lease — can send an 8-byte header declaring 0x7FFFFFFF and then stream body bytes indefinitely. Every byte is retained in RAM, none is forwarded to Xvnc, and the connection is never closed. That's an OOM of the hermes serve process from the least-privileged role in the design.
Repro (passes on d947fc83a4):
f = RfbClientFilter(lambda: False) # viewer does NOT hold the lease
f.feed(b"RFB 003.008\n\x01\x00")
f.feed(struct.pack("!BBBBI", 6, 0, 0, 0, 0x7FFFFFFF)) # ClientCutText, 2GB declared
for _ in range(64):
f.feed(b"A" * 65536)
assert len(f._buf) == 8 + 64 * 65536 # 4MB held, nothing forwarded, socket still openThe fix looks like a couple of lines, and the escape hatch already exists — feed() raising ValueError is exactly the "we cannot frame this, kill the stream" path that unknown message types use, and display.py already turns it into a _CLOSE_PROTOCOL close. Capping the declared clipboard length (real ones are kilobytes; TigerVNC's own default limit is 1MB) and raising past that reuses the existing behaviour rather than adding a new one. Worth bounding _buf overall as well, so any future variable-length type inherits the ceiling instead of needing its own check.
Happy to push the cap plus a regression test if useful.
|
^ my agent checked it based off your tweet https://x.com/Teknium/status/2098881098885595149 |
|
@teknium1 I found a Windows regression at d947fc8 that I don't see covered in the review yet. Every nonempty I also reproduced @carlotestor's parser finding on this head. A watcher without control can declare a large clipboard message and keep growing the retained buffer. My bounded probe retained 4,194,312 bytes without rejection. Their proposed message-size and buffer caps fit the existing protocol-error path. One part of @Julientalbot's takeover finding still looks incomplete: the epoch check discards results, but acquisition acknowledges human control without waiting for an already-running browser/computer-use action to stop or finish. That leaves room for a delayed click or typing operation after takeover. This last point is from source inspection, not a live Xfce reproduction. Could we cover that overlap and confirm human control only once existing actions can no longer affect the screen? |
iowahawkeyedave
left a comment
There was a problem hiding this comment.
Test pass — independent probes + suite run (macOS host, Linux CI lane untouched)
Ran the PR's own bot-desktop suites and then wrote 16 independent adversarial probes (probe_108914.py) attacking the lease and the RFB filter from angles I didn't see covered in the existing tests. All 16 pass on d947fc8. No regressions, no new findings that block.
What ran
| Check | Result |
|---|---|
tests/tools/test_bot_desktop_{lease,runtime,install,browser,browser_fence,launcher_seed}.py + test_display_ws_ticket.py + test_display_ws_drop_keeps_lease.py + test_display_methods.py (via scripts/run_tests.sh, Python 3.11) |
17 passed, 0 failed, 2 skipped (linux_only — correct on darwin, they run on the Linux lane) |
vitest: screen-connection.test.ts, sibling-ws-url.test.ts |
4 passed |
tsc --noEmit |
clean |
| 16 independent probes (below) | 16/16 pass |
Probe results (the ones worth naming)
RFB filter (rfb_filter.py):
- Byte-by-byte chunking across the handshake boundary with input gated — handshake forwards, KeyEvent fully dropped.
- Lease flip mid-stream: the very next key event from the new holder reaches Xvnc; gated viewers get nothing.
- QEMU Extended KeyEvent (255) gated identically to KeyEvent — the noVNC pseudo-encoding switch doesn't open a hole.
- Split input message across two
feed()calls with the gate flipping between halves: the message is dropped (gate checked at message completion). Fail-closed direction — a message can't sneak through half-authorized. I initially expected gate-at-start semantics; the implemented behavior is the safer one. - Unknown message type 77:
ValueError, stream killed — no blind forwarding, no hiding input behind an unframed type. - ClientCutText with negative (TigerVNC extended clipboard) length: framed correctly and gated for non-holders.
- FramebufferUpdateRequest / SetEncodings still flow for watchers — view-only viewers actually get a picture.
Lease (lease.py):
- Corrupt
lease.json→ holder=human (fail closed). Verified, matches the docstring. - Deleted
lease.json→ holder=agent. This is the one fail-open path — defensible (deleting the file requires the same local access as everything else the agent can already do) but worth a line in the docs next to the "not a security boundary" note. - Stale viewer
release("viewer-STALE")whileviewer-Aholds: lease untouched, epoch unchanged. A closing window can't yank control from the person who took over after it. - Cross-process fence: a fresh subprocess reading the same
HERMES_HOMEsees the parent'sacquire— the fcntl-on-disk design holds across process boundaries, which is the whole point. assert_agent_may_act()raisesHumanHasControlwhile a human holds; epoch strictly increments per transition;wait_for_releaseunblocked within one poll tick of a hand-back (0.41s).
Docs honesty check: bot-screen.md states plainly that screens are "work surfaces, not security boundaries" and bots share the host account — good. That's the right place to also note the deleted-lease fail-open when you next touch the docs.
Verdict
The control guarantees survived a genuinely adversarial pass. The two prior review rounds plus this one have hit it from lease semantics, transport framing, process boundaries, and the install/sudo path. Ship it.
Consolidated review of
|
iamlukethedev
left a comment
There was a problem hiding this comment.
Review of d947fc83a461
Thanks for the thorough prior rounds — the on-disk lease, server-side RFB gate, epoch discard, abnormal-disconnect exclusion, and connection-scoped pane/install listeners are in good shape. I’m still blocked on a few control-boundary and regression gaps that show up on this head.
Request changes
1. Windows computer use broken by unconditional fcntl (P1)
Every nonempty computer_use action imports handoff → bot_desktop.lease, which does import fcntl at module import (lease.py). That raises on native Windows before backend selection, even with Bot Screen disabled / unsupported. display.status hits the same import via _display_snapshot. Please keep Linux lease integration behind the supported-platform path (or a portable lock), and add a regression that handle_computer_use({'action': 'capture'}) still works when fcntl is unavailable.
2. Unbounded RFB ClientCutText buffer (P1) — agrees with @carlotestor
RfbClientFilter._message_length trusts the client-declared length and feed() grows _buf with no cap. A watcher who never held the lease can declare a huge clipboard message and retain megabytes (or worse) in the hermes serve process. Cap declared clipboard length (TigerVNC’s ~1MB default is a reasonable ceiling) and overall _buf; raising ValueError already maps to _CLOSE_PROTOCOL in display.py.
3. Local real-profile CDP sessions bypass the browser lease fence (P1)
_shares_bot_desktop_browser returns false whenever cdp_url is set, but real-profile local sessions always use CDP and this PR launches their headed Chromium with the Bot Desktop env. Under real-profile + headed + Bot Desktop, clicks/snapshots can still hit the screen the human is using. Please fence by local launch provenance; keep the exemption for unrelated remote / user-attached CDP.
4. Epoch check is after image persistence / aux-vision (P1)
In handle_computer_use, the post-dispatch epoch comparison runs after _dispatch → _capture_response → _capture_view, which may already persist the screenshot and route through aux vision. Returning human_has_control cannot retract that. Validate the admission epoch before persistence and vision routing (keep the final fence as belt-and-braces). Same idea for capture_after / spilled accessibility data.
5. Takeover vs in-flight input (P1 / clarify) — agrees with @helix4u
Discarding the result after an epoch change is necessary but not sufficient if an already-running click/type can still deliver to the seat after acquire. Prefer acknowledging human control only once admitted actions can no longer affect the screen, or cancel/fence at the backend boundary. Happy to treat a clear design note + test as enough if full cancel is hard.
Also please fix (P2)
- Portal listener still matches
profile_keyonly (screen-portal.tsx); pane/install useisEventForBotScreen. Two hosts with the same~/.hermespath can still cross-update the portal cache — finish the connection-identity fix there and cover it inscreen-connection/ portal tests. - Closing the Screen pane calls
WebSocket.close()without a status (often observed as1005); the bridge only releases on1000/1001, so Hand back via “close the pane” can leaveholder=humanand the bot paused. Explicit lease release or a deliberate clean close on intentional teardown; keep exclusion on real drops (1006). - CLI
hermes computer-use screen stoponly callsruntime.stop()while help text says it hands control back; gatewaydisplay.stopalready releases. Align CLI with the RPC. - Display allocation: host lock is released before Xvnc creates the X lock — two cold starts can pick the same
:N. Hold ownership until claimed / record pending reservations. - Thumbnail: mutating process-wide
XAUTHORITYis racy across concurrent profile thumbnails in worker threads — isolate (e.g. subprocess or thread-local capture). - Lease reader hardening:
{}/{"holder":"invalid"}currently admit the agent;[]raisesAttributeError(not caught). Fail closed on unknown/invalid shapes.
Nits
- package-lock
@novnc/novnc@1.7.0looks intentional; please addci-reviewedonce the bootstrap-installer lockfile bump is confirmed expected. - Would love regressions for: Windows import gate, RFB size caps, real-profile fence, portal connection match, pane-close release, CLI stop release.
CI
Required checks on this head look green (All required checks pass). Desktop E2E was skipped. Mergeability is still blocked pending review/labels.
Happy to re-review quickly once the P1s are in.
Xipong
left a comment
There was a problem hiding this comment.
Reviewed d947fc83a461f098066c71ee6e28161736b87f53. I recommend fixing the following before merge. Three additional defects and one residual occurrence of an existing finding; no style requests.
[P1] Ordinary Windows computer_use now imports Unix-only fcntl
Locations: tools/computer_use/tool.py::handle_computer_use, tools/computer_use/handoff.py, tools/bot_desktop/lease.py.
Every computer_use invocation now imports handoff before dispatch; handoff imports the lease, whose module-level import fcntl is unconditional. Native Windows Python does not provide fcntl. Even an ordinary Windows capture/click with Bot Desktop disabled therefore fails before reaching the CUA backend. runtime.start's Linux guard is too late.
Gate the Bot Desktop integration before importing Unix-only code, or make the lease importable on unsupported platforms without breaking the existing Windows backend. Test ordinary computer_use with fcntl unavailable and Bot Desktop disabled. Source-traced portability regression; no native Windows execution claimed.
[P2] Clean pane closure can strand the human lease
Locations: apps/desktop/src/plugins/hermes-bots/screen-pane.tsx::detach; hermes_cli/web_routers/display.py::_CLEAN_CLOSE / _bridge.
The server releases control only for 1000/1001, but pane cleanup calls rfb.disconnect() and socket.close() without a close status. The pinned noVNC 1.7.0 Websock.close also calls the underlying WebSocket.close with no argument. A clean close with no status is reported as 1005. Closing the pane while holding control can therefore leave the persisted human lease held, blocking computer_use and local browser tools.
Actually executed protocol probe (Node WebSocket -> Python websockets server): ws.close() yielded server code 1005 and client wasClean=true; ws.close(1000) yielded server code 1000. This PR's whitelist releases only the latter. This is not an Electron/noVNC/Xvnc end-to-end test.
Explicitly distinguish intentional pane closure from reconnect/transport failure and return control on the intentional-close path, or send the intended status on that path. Do not unconditionally release in detach: attach calls it too, and abnormal-disconnect exclusion must remain intact.
[P2] The allocator releases its lock before reserving the display
Location: tools/bot_desktop/runtime.py::_allocate_display / start.
The host-wide flock ends when _allocate_display returns. start only records the selection in the current profile, then spawns Xvnc. A second profile/process can enter before Xvnc creates /tmp/.X20-lock and select the same :20; the allocator does not inspect other profiles' selections or keep pending reservations. At least one desktop can fail despite other display numbers being free. Separate CLI processes can hit this regardless of per-gateway RPC serialization.
An isolated probe with the extracted allocator, real flock and temporary profile directories returned :20 for both profiles after recording A's local selection but before Xvnc startup. No live Xvnc concurrency test claimed.
Hold the allocation lock until the server claims its display, or create a host-wide pending reservation with failure cleanup. Test two starts interleaved before Xvnc readiness.
[P2, residual of the existing review] The portal still matches lease events by path alone
Location: apps/desktop/src/plugins/hermes-bots/screen-portal.tsx::useScreenPortalState, display.lease listener.
isEventForBotScreen was added to the pane/install listeners, but this listener still checks only payload.profile_key === profileKey and writes the shared screen store. A host-B event can still overwrite host A's control state when both homes have the same path. Since VIEWER_ID is window-wide, A can claim this viewer is in control while A's backend remains agent-controlled.
This is the unconverted occurrence already identified in @Julientalbot's portal review, not a new independent discovery. Apply the connection/profile predicate here too and test the actual hook with two hosts sharing a path; testing the predicate alone does not detect an unconverted consumer.
Scope: full diff and related paths inspected, with the protocol/source-isolated probes above. No full repository suite or live cross-platform desktop run claimed.
Xipong
left a comment
There was a problem hiding this comment.
Follow-up review of d947fc83a461f098066c71ee6e28161736b87f53, additional to review #5188188627. Four further findings below. I did not repeat the Windows import, close-code, allocator or cross-host portal findings.
[P1] Locally launched real-profile Chrome is incorrectly exempted from the human lease
Location: tools/browser_tool_session.py::_shares_bot_desktop_browser, together with _create_local_session and tools/browser_tool_real_profile.py::_launch_real_profile_chrome.
cdp_url does not establish that a browser is external to Bot Screen. With browser.use_real_profile and browser.headed enabled, Hermes launches the local profile-copy Chrome with the Bot Desktop environment introduced by this PR. _create_local_session records that browser as features={local: true, real_profile: true} and gives it a loopback cdp_url. The new predicate immediately returns false for any CDP URL, so actions and observations against this browser skip both the pre-dispatch human check and the epoch fence, even while a person is using its window on Bot Screen.
Controlled probe: actual _create_local_session, _run_browser_command and _shares_bot_desktop_browser function excerpts, a real persisted HUMAN lease, and stubbed CDP resolution/final command execution. The local real-profile snapshot reached the execution seam and returned a synthetic private-page marker. The ordinary no-CDP local session was refused with human_has_control as a positive control. No Chrome process or real page was used.
Please classify the browser by its owned local launch/display provenance rather than treating CDP transport as proof of a remote browser. Keep genuinely unrelated cloud/user-CDP endpoints exempt. Cover this supported real-profile configuration, not only {session_name: ...} sessions.
[P2] Starting Bot Screen does not rebind an already cached CUA backend
Locations: tools/computer_use/cua_backend.py::cua_driver_child_env, tools/computer_use/tool.py::_get_backend, and tools/computer_use/cua_backend_session.py::_lifecycle_coro.
The new display environment is applied only when a driver transport is spawned. _get_backend reuses the existing backend whenever the permission mode matches; neither display.start nor the tool-boundary hook invalidates it when the profile's display changes. The MCP session is long-lived. Consequently, using computer_use before Start Screen, then starting the screen and continuing in the same session, can leave the driver bound to the original display instead of the screen being viewed. On a Linux workstation that can be the original human seat; a pre-existing headless driver can likewise remain attached to its old environment. A changed display/Xauthority after restart needs the same treatment.
Controlled cache probe using the unchanged _get_backend/_install_backend excerpts and an environment-recording backend: after changing the available environment from :0 to :20, the next call returned the identical backend with its original :0 spawn environment; only one start occurred. The real MCP code was traced to confirm that its child environment is sampled at transport creation, not each action. This is not a live CUA test.
Please include the effective profile/display identity and runtime incarnation in backend reuse, and replace the transport/clear sticky target state before dispatching on a different screen. Add a same-session before-Start/after-Start and screen-restart test.
[P2] Handoff requests from another process never reach the open screen's event listener
Locations: tools/bot_desktop/lease.py::on_change, tui_gateway/methods_display.py::_install_lease_listener, tools/computer_use/handoff.py, and the Screen pane/portal status subscriptions.
The lease is correctly shared on disk, but notification is still process-local. The display gateway broadcasts only transitions delivered to its own lease.on_change listener. A messaging gateway, CLI or isolated worker can call request_handoff in another process, update lease.json, and wait for hand-back without emitting any event in the process serving the Desktop viewer. An already-open pane only fetches status initially/reconnects and otherwise follows those events; thumbnail refreshes carry no lease state. Thus the advertised Bot needs you indicator/reason need not appear until the user manually refreshes or reopens the screen. This finding concerns the screen notification, not whether the model separately tells the user in chat.
Executed two-process probe with the byte-identical lease.py blob 2a3b84d5e5d39f57853682b4a9af741897493ae2, real files and flock. The child request was visible to get() in the parent, but the parent's subscriber received zero notifications; a same-process transition delivered one. wait_for_release remained blocked on the pending request. Only the home/key helper was shimmed to a temporary directory.
Please bridge file/epoch changes into the display server's event stream, or provide bounded status reconciliation that includes handoff requests. Test an already-open viewer with a request emitted by another process.
[P2] The final exec drops the launcher's Xvnc cleanup trap
Location: tools/bot_desktop/launcher.sh: trap ... EXIT followed by exec dbus-run-session ....
Xvnc is a background child of the shell, but the shell is then replaced with dbus-run-session. A successful exec does not run or preserve that shell EXIT trap. If the Xfce panel/private session exits independently of runtime.stop, Xvnc is therefore not cleaned up. Once the launcher PID is gone, _launcher_pid reports no running launcher and stop() returns without signalling the orphaned group. A later start can allocate another display and unlink the old RFB pathname while the old X server continues to exist. Explicit killpg-based Stop working does not cover this independent-session-exit path.
Executed the exact shell lifecycle pattern with sleep standing in for Xvnc and /bin/true standing in for the session: the session/launcher exited, while its server child remained alive. The probe cleaned that child up. This demonstrates the exec/trap issue, not a live Xfce crash.
Keep a supervisor shell/process alive to wait for and reap both children, with cleanup on session failure as well as explicit Stop. Test exit of the session child and verify the X server, socket and display lock are gone.
Validation scope: source tracing plus the controlled probes explicitly described above. No full repository suite, real Chrome/CUA/Xvnc desktop, or Electron acceptance run. A suspected noVNC disconnect/reconnect issue was investigated and excluded: its actual clean-disconnect path triggers the pane's existing reattach effect.
erosika
left a comment
There was a problem hiding this comment.
review of d947fc83a461
@teknium1 requesting changes. the RFB filter itself is solid, agree with @carlotestor and @iowahawkeyedave there. the problems are around it: who the server thinks "the viewer" is, what a display ticket lets you do, and a couple of startup races. most of the below isn't in the earlier reviews yet.
what I ran
- PR suites via
scripts/run_tests.sh: 17 passed, 2 skipped (linux-only, I'm on macOS) tests/tools+tests/tui_gateway+tests/hermes_cli: 8045 passed, 11 failed. 10 fail the same way on main. the other one (test_doctor_structural_corruption) is flaky under load and passes 3/3 on its own. no regressions from this PR- desktop:
tscclean, eslint clean, the two new vitest files pass - no live Xvnc/Xfce run, no linux box on hand
confirmed from earlier reviews
| issue | who found it | note |
|---|---|---|
windows computer_use breaks on import fcntl |
@helix4u, @MrD1az, @iamlukethedev | confirmed. check-windows-footguns passed, so the guard doesn't cover this |
| closing the pane doesn't hand control back | @MrD1az | reproduced separately. close() with no code arrives as 1005, bridge only releases on 1000/1001 |
| sidebar Screen row ignores which host an event came from | @MrD1az | screen-portal.tsx:105 |
| clipboard message can grow memory without limit | @carlotestor | |
| real-profile browser skips the lease check | @MrD1az | |
| takeover drops results but doesn't stop actions already running | @helix4u | also: a key held down when control flips never gets its key-up, so after hand-back the agent's next type can go out with a stuck modifier |
| screenshot is saved / sent to vision before the takeover check | @MrD1az | |
display number race, CLI stop not releasing, thumbnail XAUTHORITY swap, bad lease file shapes |
@MrD1az, @iamlukethedev | agree with all |
tickets and viewer ids let other clients in
a display ticket also logs you into the main gateway socket
display.observe mints a ticket meant only for /api/display/ws. but /api/ws accepts any ticket and never checks what it was minted for (web_server_chat.py:273), only the display route does (display.py:45). so a display ticket works as a gateway login, with a user id the client picked. that includes the embedded-TUI internal credential, which is otherwise not treated as a real login.
fix: reject provider == "bot-desktop" tickets at /api/ws, or keep them in a separate store.
anyone who knows the holder's viewer id can type into the screen
- the client picks its own
viewer_id(methods_display.py:105,161) - the bridge lets input through when that id matches the lease holder (
lease.py:197-199) - every lease change is broadcast to all clients with the holder's id in it (
methods_display.py:44-45)
so another client reads the id off the broadcast, opens a stream with it, and gets keyboard + mouse next to the real holder. nobody is evicted and the chip doesn't change. if every authenticated client is meant to be equally trusted this is a docs line. if not, the server should assign the id, tie it to the connection, and leave it out of the broadcast.
desktop startup race
starting the same profile twice at once orphans the first desktop
runtime.py:249 only counts the desktop as running once the env file exists, and the launcher writes that after Xfce is up (~5s). nothing locks around the check. a second start in that window (Start button + auto-start, or two clients):
- deletes
envand launches anotherlauncher.sh - that launcher deletes the live
rfb.sock(launcher.sh:46) and writes a new X cookie (launcher.sh:52) - its Xvnc fails, but it already overwrote
launcher.pid(runtime.py:280)
end state: first desktop still running with no socket, a changed cookie, no pid on record, and stop() can't find it. this is a different bug from the two-profile display race. fix: one per-profile lock across start and stop.
Screen UI in Hermes Desktop
sidebar Screen row remounts on every render
gateway-groups.tsx:346 passes a new inline function to ContribRender, which uses it as the component type. new type every render → unmount + remount, and while status is unknown each remount sends another display.status. other ContribRender callers pass a stable function.
a profile group with no connection can open another host's bot
screen-portal.tsx:166 treats connectionId === null as a match for any connection. a local group called research can open, and take over, a remote bot also called research.
older backends show a Screen UI that never settles
if display.status fails the state stays unknown, and only unsupported hides the UI. on a backend without display.*, every profile group shows an empty Screen row and the Scheduled Jobs preview says "Checking the screen…" forever.
"control-taken" message never shows
noVNC's disconnect event only has { clean }, no reason, so reason.includes('control-taken') in screen-pane.tsx:137-145 is never true. eviction lands on the clean → re-attach path by accident.
X cookie on argv, lease reads on the event loop, install + UI loose ends
launcher.sh:52puts the X cookie on thexauthcommand line, briefly readable by other local users.xauth source -over stdin avoids itdisplay.py:100readslease.jsonsynchronously on the event loop for every client message, so every mouse move- install timeout kills
sudo, not the package manager under it - the install sudo card has no session id, so it attaches to whatever chat is open, marks that chat as needing input, and can go stale if you switch chats
- the stale-preview check only counts errors, so a hung socket takes minutes to show "Last seen"
screen.rechecki18n key is unused.resolveSiblingWsUrlcopiesvoice-playback.tsinstead of replacing it. config comment sayshermes desktop, the command ishermes computer-use screen
same-user processes can still capture the login
agree with @Julientalbot that one shared OS user isn't a security boundary. one case worth spelling out in bot-screen.md: a same-user process can swap bot-desktop/rfb.sock for its own relay, the bridge will dial it (display.py:69-77), and it captures what the human types during the login the takeover exists for. "fail closed" in the lease.py docstring reads stronger than that.
regression tests to add with the fixes
- closing the pane releases the lease (the real
close()path, not a faked 1000) /api/wsrejects display tickets- a foreign
viewer_idgets no input - two starts on the same profile at once
- a group with no connection doesn't match a remote bot
- windows:
computer_useworks withoutfcntl
Independent review round — Hermes agent fleet (6 workers), native Windows + executed probes, at
|
Homer review (native Windows 11 + packaged Desktop)Tested Host: Windows 11, New evidenceP1 confirmed on native Windows (was UNKNOWN): Chain: Same crash on:
This matches @helix4u / @iamlukethedev / @Xipong / @MrD1az F3. Their probes
Then PR tests on native Windows (
Desktop viewer (Windows packaged app → Linux
Pane close / 1005: not exercised (no Open Screen in this Desktop build). Not re-reportedRFB clipboard cap, real-profile CDP fence, epoch-before-vision, in-flight AskPlease gate |
BearHuddleston
left a comment
There was a problem hiding this comment.
Reviewed the exact head d947fc83a461f098066c71ee6e28161736b87f53. Read the existing top-level discussion, formal reviews and inline threads before investigating, then refreshed them after both review workers finished and validated their findings.
Three inline findings remain after deduplication: two additional Desktop lifecycle/identity defects and real execution of the previously UNKNOWN human-first dock-browser handoff. I am not re-filing the Windows import, lease/fencing, RFB buffer, portal connection, close-code, allocator, sessionless sudo, or other already-covered issues.
Verification:
- The nine focused Bot Desktop/display Python files pass via
scripts/run_tests.sh: 19 passed, 0 failed. - Executed React/TypeScript component, gateway-pool and event-bus probes: reproduced the missing route retention and cross-bot thumbnail reuse. Retaining the route was a passing positive control. The sessionless sudo probe also reproduced an existing report and is excluded from the inline findings.
- Real agent-browser 0.26.0 and Chrome for Testing 151.0.7922.34 on isolated Xvfb: the human-first handoff assertion fails; agent-first and same-session recovery after closing only the human-launched browser succeed. Temporary HOME/HERMES_HOME and user/network namespaces; no external page or credentials.
- Additional corroboration of Xipong's existing review, not new findings: real cua-driver 0.26.1 stays on the original 640×480 display after publishing an 800×600 Bot Screen, while a fresh session captures 800×600; the real launcher with extracted Xvnc and a controlled panel exit leaves a responsive X server after
runtime.stop()returns false. These external behavior-contract probes intentionally fail on this head; they are not passing regressions.
Limits: no full Electron/noVNC/Xfce acceptance run or full repository suite. Desktop probes used transport/UI-boundary doubles; browser setup/discovery was controlled, but the Chromium/agent-browser execution was real. Exact-head CI reports required checks passing, with Desktop E2E skipped. No tracked source edits, commits, pushes, host package installation, or live-profile changes.
| from tools.bot_desktop.browser import dock_launch | ||
| if (browser := dock_launch()) is not None: | ||
| # first-run / default-browser dialogs would sit between the human and the bot's tabs | ||
| child_env["HERMES_BD_BROWSER_EXEC"] = f"{browser[0]} --user-data-dir={browser[1]} --no-first-run --no-default-browser-check" |
There was a problem hiding this comment.
[P2] Make the human-first dock browser attachable by automation
This launches raw Chromium with the shared profile but no automation endpoint. On a fresh bot, Take over → open the dock's Browser before any agent-browser call → Hand back leaves that Chromium owning the profile singleton. The subsequent agent-browser launch exits through the singleton without obtaining CDP, so the bot cannot continue browsing while the human-opened instance remains running.
I executed this dock argv and the real _run_browser_command/environment/process path using agent-browser 0.26.0 and Chrome 151.0.7922.34 on isolated Xvfb. The result was success:false, Chrome exited early (exit code: 0) without writing DevToolsActivePort. The original browser stayed alive and had no DevTools port file. Closing only that browser made the next call in the same profile and agent session succeed; an independent agent-first control also succeeded. No live login or UI acceptance run is claimed. This supplies the real execution missing from the earlier UNKNOWN singleton concern.
Please have the dock open an automation-owned browser, or implement an explicit shared-browser attachment lifecycle. A common binary/user-data-dir alone does not establish an automation connection; cover both launch orders.
There was a problem hiding this comment.
Fixed in d732fad0af00 (rebased into d65af42). The dock launcher now runs Chromium with --remote-debugging-port=0; tools/bot_desktop/browser.py::running_instance_cdp_port trusts DevToolsActivePort only when the SingletonLock pid is alive and the port accepts a connect (and excludes the instance the calling agent-browser session launched itself), and the local argv gains --cdp <port> when that instance exists. Live on this host, both orders: human opens the dock browser first → browser_navigate succeeds in the same single Chromium process; agent first → dock click joins the agent instance and later navigates still succeed. Before the fix the human-first order failed exactly as you described (exit 21, no DevToolsActivePort).
|
|
||
| try { | ||
| await displayRequest(bot, 'display.install') |
There was a problem hiding this comment.
[P2] Retain the gateway route while Screen waits for events
Open an inactive local bot's Screen directly without activating its chat. This RPC uses a pooled secondary, but neither the pane nor this card holds host.retainProfile. requestGatewayForAgent releases its request lease in finally and disposes an otherwise unowned socket as soon as display.install acknowledges. The actual install runs in a background worker, so subsequent sudo/log/done events have lost their receiving transport: the card stays Installing with its button disabled. display.observe has the same lifetime problem for subsequent lease events even while the separate RFB socket remains open.
A probe through the actual install component, display routing, gateway pool and event bus (transport doubled) reproduced closure at ACK, loss of the completion event and the stuck button. Holding the existing retainGatewayForAgent('local', 'worker') as a positive control delivered completion and re-enabled the button.
Acquire host.retainProfile(botScreenRoute(bot)) before starting this event-driven operation and keep it for the appropriate Screen/install lifetime, releasing on cleanup. This is separate from the already-reported sudo prompt's chat association and the corrected worker-context propagation.
There was a problem hiding this comment.
Fixed in 1eed9c82e27d (rebased into d65af42). retainBotScreen(bot) feature-detects host.retainProfile the way group-turns.ts does; the install card retains before display.install and releases on done/failed/unmount, and the pane holds a retention for the attach lifetime (acquired before display.observe, released on detach/unmount). Test: screen-install-retention.test.tsx asserts retain-before-request and release-on-done.
| useEffect(() => { | ||
| if (!running) { | ||
| setDataUrl(null) | ||
| setMisses(0) | ||
|
|
||
| return | ||
| } |
There was a problem hiding this comment.
[P2] Clear the thumbnail when its owning bot changes
RoutinesPane renders this hero without an owner key (cron.tsx:1245), so selecting another bot reuses the hook state. If both bots have cached running status, running stays true and this branch never clears dataUrl. The newly selected bot B consequently shows bot A's pixels beneath B's current status/CTA until B's first thumbnail succeeds. A slow or failing B can keep displaying the wrong bot's screen, not merely an old frame from B.
I reran a probe rendering the actual ScreenHero: after A's thumbnail resolves, rerender with B and defer B's response. The image remains A's and the caption still says Live; resolving B finally replaces it. Thumbnail transport and portal tone were controlled; the unkeyed owner-change caller was source-traced.
Reset image/failure state on the connection+profile identity change, or key the hero by that identity. Add a switch-between-two-running-bots test with B's response pending/rejected. This is distinct from the earlier same-bot failed-refresh staleness report and the portal's wrong-host event filtering.
There was a problem hiding this comment.
Fixed by @whyyagswhy's 7bc75c0 (cherry-picked with authorship, rebased into d65af42): ScreenHero is keyed by botSelectionKey(bot) so dataUrl/misses/stale reset on an owner change, and a late thumbnail from the previous bot is discarded. Test: screen-isolation.test.tsx "hero reset on bot switch" + "late-thumbnail discard".
|
@teknium1 trying to test this on a Nous-hosted Hermes instance, with Hermes Desktop on Windows as the viewer. What is the supported way to get the matching gateway preview onto the managed Cloud instance? The Desktop side has been updated to this PR at Running the documented The installed host also lacks I understand the PR docs explicitly support Hermes Cloud connections, and Install on host installs the Linux desktop packages once the gateway has the new handlers. The missing part here appears to be getting that gateway code onto the hosted instance, not running Linux on my Windows viewer. Can Nous enable a matching preview deployment, or is there a supported self-service branch/preview option I have missed? Happy to test the actual screen takeover, manual login, hand-back, and continued Cloud operation with the PC off. Related to #92524; posting here rather than opening a duplicate. |
… instead of silently ignoring them profile_dir() accepted only an absolute AGENT_BROWSER_PROFILE; `~/pin` and `pin` fell back to the default user-data-dir while the docs said setting the variable pins your own. The dock's Browser icon and agent-browser both derive their identity from profile_dir(), so the silent fallback was at least confusing and, combined with any other consumer of the raw value, a way to end up on two jars. `~` is expanded and a relative path is anchored at the profile's HERMES_HOME (the same root runtime.state_dir() uses), so per-profile pins stay per-profile. The docstring and the bot-screen docs state the resolution rule. Contributor PR #112708 (@Rook-CodeVolt) proposed documenting "absolute only"; this takes the behavioural fix instead. Closes #110029
…opping the viewer's stream launcher.sh starts Xvnc with -AcceptSetDesktopSize so the viewer can fit the screen to its pane, but rfb_filter.py had no frame for client message type 251 and raised "unknown RFB client message type 251", closing the WebSocket on the first resize request. SetDesktopSize (u8 type, pad, u16 width, u16 height, u8 nScreens, pad, then 16 bytes per screen) is framed like the other variable-length client messages and forwarded intact; it carries no input, so watchers without the lease send it too. A truncated message waits for the remaining bytes. Closes #110039
…creen like computer_use does bot_desktop.auto_start had one trigger, computer_use dispatch. The browser tool built its headed child env from desktop_env(), which only reads the env the launcher already published, so on a fresh headless profile with browser.headed and auto_start both on the first browser use got no DISPLAY and nothing said why. _build_browser_env() now calls the same ensure_started_for_tool() hook when a headed browser is requested (headless browsing never brings a screen up), with the hook's own failure handling — a failed start falls through to the tool's usual "no display" diagnosis, exactly as for computer_use. Docs updated: "first use" covers the first computer_use call or headed browser use. Closes #110050
…creen; the env builder stays pure _build_browser_env() is shared by the npx cache warmer (hermes update / doctor --fix), the lazy Chromium auto-installer, the Lightpanda engine's env and browser_use_cli. Hooking bot_desktop.auto_start there made every one of those block up to 15 s spawning Xvnc+Xfce with browser.headed on, against desktop_env()'s own "never starts anything" contract. The hook now sits where a headed Chromium is actually spawned for a tool action, mirroring computer_use dispatch: _spawn_and_collect (the first agent-browser command forks the daemon; Lightpanda engine excluded), the Chrome fallback from Lightpanda, and the real-profile Chrome launch. The regression test asserts both halves: the env builder never starts the screen, the headed Chromium spawn does, a headless or Lightpanda spawn does not.
… to Xvnc Framing client message 251 (so the stream no longer dies) is kept, but the message resizes the bot's framebuffer under a working agent — Xvnc runs -AcceptSetDesktopSize — so forwarding it from a viewer that does not hold the lease was a fail-open flip for a server-mutating message. The shipped pane sets resizeSession=false anyway; SetDesktopSize now sits in _INPUT_TYPES and only the lease holder's resize reaches the server. Test covers both legs.
… the socket it binds _reap_orphaned_server only knew the orphan via /tmp/.X<n>-lock, so a launcher SIGKILLed while a /tmp cleaner also took the lock (case B of #109941) — or a failed start() that dropped <sd>/display — left one live Xvnc leaked per occurrence; the allocator merely moved to the next number. When the lock names nothing alive, the reaper now scans for the Xvnc whose command line binds this profile's rfb.sock (the same ownership proof it already required), and skips the lock unlink when no number is recorded.
…argument the base branch added
…live The janitor reaped the bot's headed Chromium after 120 s of AGENT inactivity — which is exactly the state a human takeover (login, 2FA) puts the agent in — so the browser died under the human mid-login. The janitor now counts a human-held lease as activity for the browser the human shares with the bot, and the agent-browser daemon's own idle timer (which cannot see the lease) steps back on the Bot Desktop so the lease-aware janitor owns that browser's lifetime; a crashed hermes still leaves it to the orphan reaper. Fixes #110064
…he threat model (#110040)
…ot Desktop display
…-literal guard Main now rejects literal /tmp paths in favour of the scratch-dir resolver. The X11 protocol fixes .X<n>-lock and .X11-unix/X<n> under /tmp; these lines detect or document that location, they do not pick scratch space.
…top (opt-in) Watching the bot drive its screen is the point of Bot Screen, but until now you learned it had happened only afterwards. With "Open Screen when the bot uses it" checked on a bot's row menu, the first live tool.start for a screen tool (computer_use, browser_*) on a session the bot owns brings its Screen tab forward. Why opt-in and fenced: Desktop's rule is offer, don't hijack. The raise is per bot (BotMeta.screenAutoOpen, rides profile ui_meta like pin/hide), never moves keyboard focus (openBotScreen reveals the pane; the viewer grabs keys only on Take over), never fires for replayed history (the reconnect replay re-dispatches parked frames, so the wake is rate-limited to one per bot per 30 s rather than trusting seq), does nothing while the tab is open, and a manual Close mid-run holds until the run has been quiet for a cooldown. Live (isolated headless Electron + real serve backend, CDP): opt-in off → tool.start opens nothing; toggled via the real menu → toast + aria-checked true → tool.start opens "Hermes · Screen" (Screen is off state), focus stays on the row; second call is a no-op; real Close → held at +2 s and +29 s, raised again at +31 s. Rule borrowed from thomasbek3/hermes-bot-kit computer-viewer's auto-connect.
…it (#115184); server registration, i18n, contracts
…lution ratchet Main's hermes_platform ratchet (tests/test_managed_runtime_resolution.py) flags every bare shutil.which outside hermes_platform/. Bot Screen's four sites probe Xvnc, xfce4-session, a system Chromium and apt/dnf/pacman on the gateway host, which is the control host the resolver describes; none of them installs or starts anything and the install path runs only after the user's sudo approval.
A bot on a headless Linux gateway gets its own Xfce desktop that Hermes Desktop streams live; you can take over to log in / solve 2FA, then hand the screen back and the bot continues with your session.
Related: #92524 (the "let me open the bot's screen, log in myself and hand it back while the agent keeps running after my PC is off" request; this PR is the Desktop + Linux-host half — a bot's own screen, takeover, hand-back and a shared browser profile. The hosted/cloud-browser leg and dashboard integration are not in it, so it does not close the issue). Also related: #108592, #97859 (earlier Bot Desktop direction), #17258 (Docker Xfce), #90380 (CUA backend seam), #90374 (agent acting on the human's seat).
What changed
tools/bot_desktop/—launcher.shstarts TigerVNCXvnc(Unix socket, 0600,-SecurityTypes None, no TCP) plus xfwm4 / xfce4-panel / xfdesktop / xfsettingsd under a privatedbus-run-session, one per profile (<HERMES_HOME>/bot-desktop/).runtime.py= start/stop/status, free display allocation, publishedDISPLAY/XAUTHORITY/DBUSenv.lease.py= who drives the screen (agent | one human viewer), per profile.rfb_filter.py= stateful RFB client-stream parser that drops KeyEvent / PointerEvent / ClientCutText / QEMU Extended KeyEvent from viewers that do not hold the lease (server-side view-only, not noVNC'sviewOnlyhint).hermes_cli/web_routers/display.py—/api/display/ws: raw RFB over WebSocket, authenticated with a one-shot 30 s ticket minted over the already-authenticated/api/ws; lease flip closes evicted viewers with4000 control-taken.tui_gateway/methods_display.py—display.status|start|stop|observe|lease.acquire|lease.release+display.leasebroadcast event.tools/computer_use— every action (capture included) is refused withcode: human_has_controlwhile a human holds the lease (takeover is always human-initiated: when the bot hits a login/2FA/CAPTCHA it says so in its reply and ends its turn; the person takes over from the pane, does the step, hands back and tells it to continue; there are no agent-side handoff actions); cua-driver and headed Chromium inherit the bot's display so the agent never acts on a seat the human is sitting at. Auto-start is opt-in (bot_desktop.auto_start, default off): the Screen pane's Start button is the normal path; with the flag on, a headless host with the packages present starts the screen at the tool boundary on first use.display.thumbnail, one JPEG grab every 4 s while visible; read-only, never touches the lease) above the title and the routines, click to expand into live access — a compact Screen row under each gateway/profile header in the Sessions sidebar (newsidebar.gatewayGroup.headercontribution area; the plugin owns the box, core only exposes the slot), and Bots → right-click → Open Screen. The box shows Live · bot in control / Stopped / Not installed on host from the samedisplay.status+ lease events as the pane. The pane itself: noVNC (@novnc/novnc) pane with chip (Bot is in control / You are in control / Another viewer is in control), Take over / Hand back, red border while you drive; closing the pane hands control back.data-terminalon the canvas host so the app's type-to-focus shortcut can't steal keystrokes.display.install, which runs the distro package command (apt/dnf/pacman) on the gateway host as a supervised child and streamsdisplay.install.log/display.install.done. If sudo needs a password the host raisesdisplay.install.sudo.requeston the caller's own WebSocket and the renderer shows the existing masked SudoDialog (answering ondisplay.install.sudo.respond, redacted from the gateway trace likesudo.respond); an empty answer cancels without spawning the package manager; one install per profile at a time. Nothing installs onhermes update; the only triggers are this button and the CLI.tools/bot_desktop/wallpaper.png(Nous gradient) seeded as the backdrop, first dark GTK/icon theme the host ships, dark translucent panels; own panel layout with a dock of only launchers whose program exists (browser pinned to the one the bot drives, so a human who takes over lands in the bot's browser profile). Users install whatever they like on the host; it lands in the Applications menu.hermes computer-use screen status|start|stop|install(apt/dnf/pacman package lists, deliberately not thexfce4metapackage: no screensaver / power manager / polkit agent on a headless desktop).website/docs/user-guide/features/bot-screen.md, registered insidebars.ts.Live evidence (Linux host,
hermes serveheadless + Hermes Desktop via CDP, real Xvnc/Xfce)computer_use captureon a headless profile (noDISPLAY)0x0, "no DISPLAY is set"1440x9001440x900streaming the Xfce panel +hermes-bot-terminalecho HUMAN-TYPED-VIA-HERMES-DESKTOP⏎4000 control-taken, human leasesecond-laptopagent,computer_usecapture works againdefaultprofile header, same state, click opens the panesudoline to copydisplay.install.logcarries the same line; a second click while one runs is refusedTwo bugs found and fixed only by the live pass: noVNC never received
openwhen the socket was dialed before its dynamic import finished (pane stuck on "Connecting…"), and the RFB filter rejected message type 255 (QEMU Extended KeyEvent), which noVNC switches to as soon as Xvnc advertises the pseudo-encoding, so the first keypress closed the stream.Validation
tests/tools/test_bot_desktop_install.py(empty sudo answer cancels without spawning; second install per profile refused, slot released),tests/tools/test_bot_desktop_lease.py(RFB filter across byte-by-byte chunking incl. QEMU key; tool refusal while human holds),tests/hermes_cli/test_display_ws_ticket.py;apps/desktop/src/lib/sibling-ws-url.test.ts.tests/tools,tests/tui_gateway,tests/hermes_cliviascripts/run_tests.sh: only the 6 failures that also fail on pristineorigin/mainon this host (modal/parallel SDK,sortpayload, update live-system guard).tsc --noEmitclean, eslint clean, vitest 969 files / 9769 tests passed.ruff,check-windows-footguns --all,check_compat_pointers,check_subprocess_stdin: clean.tests/tools/conftest.pypins Bot Desktop binaries to "missing" so a dev host with TigerVNC installed never launches real X servers from the suite (same class as the browser-use fixture).Independent review round (reproduced → fixed → re-verified live)
An independent review of
d9525b3found five P1s and a P2; every one was reproduced, fixed in3c70635, covered by an invariant test proven red without the fix, and re-verified on a real Xvnc/Xfce screen.hermes servedid not stop a gateway/CLI process driving the same displaylease.jsonunder an fcntl lock in the profile'sbot-desktop/; every read hits the file;epochper transitionacquire→ process Bcomputer_use capture=human_has_control, no png; process Crelease→ B captures again_dispatchthat flips the lease mid-flight → result dropped,SECRETnever returned$gateway(host B could receive host A's password)SudoRequest.origin= (connection, profile) of the request; SudoDialog answers viarequestGatewayForAgenton that socket;sudo.expire/display.install.sudo.expiretear the card downgoogle-chrome→ a different user-data-dir than the bot'stools/bot_desktop/browser.py: one identity (agent-browser's Chromium +bot-desktop/browser-profile) applied to BOTH the agent env (AGENT_BROWSER_EXECUTABLE_PATH/AGENT_BROWSER_PROFILE) and the dock launcherlocalStorageon127.0.0.1:8765via agent-browser; dock click + typed URL on the screen showed BOT-WROTE-THIS in the same profile (screenshot below):20, A restarted on:21, B kept runningcopy_context()carries the HERMES_HOME override + transport into the threadScope honesty from the same round:
Closes→Related(hosted browser leg not here),auto_startdefault off,request_handoffno longer claims a Telegram/Discord message was sent (the model relays the ask in its reply).Shared browser profile: the bot wrote
BOT-WROTE-THISinto localStorage through agent-browser; the dock's Browser icon, clicked on the screen, opens the same Chromium + user-data-dir and the page reads it backSecond round (@Julientalbot, on
3c70635) — 3 P1 + 2 P2 + 2 design notes, all reproduced, fixed in7a43ce8/d947fc8_run_browser_commandbrackets every local agent-browser command with the lease (refuse + epoch fence); cloud/user-CDP sessions untouched;code: human_has_controlon the tool resultbrowser_clickandbrowser_snapshotrefused; navigate worked after hand-backhuman_holds()dropped_bridge()hermes serve:transport.abort()→ holder stayshuman;close(1000)→agent~/.hermes)isEventForBotScreen: connectionId AND profile key, one predicate for all three listenersscreen-connection.test.tsacquirekeepspending_handoffasreason; pane shows it while heldThird round (community, on
d947fc8: 10 reviews, 15 comments, 3 inline threads, sibling PR #109446) — ~40 distinct findings, every legitimate one fixed ind947fc8..d65af42Reviewers: @Julientalbot @BearHuddleston @carlotestor @iowahawkeyedave @iamlukethedev @Xipong @erosika @rahlquist @helix4u @MrD1az @Ganaderiapp @eynaudg @lEWFkRAD @dresraz @whyyagswhy @zfifteen @coe0718 @thomasbek3. Four commits by @whyyagswhy (PR #109446) are cherry-picked with authorship.
lease.pyimportedfcntlat module level; everycomputer_usecall,display.status,screen statusbroke on native Windowsfcntl(no-op lock where absent, file semantics kept);check-windows-footguns.pynow flags module-level POSIX-only importscdp_url) escaped the lease fence; a dead Xvnc with a stranded human lease unfenced itfeatures.local), armed by live DISPLAY or human leasebackend.capture(), before persist/spill/vision; alsocapture_afterdisplay.thumbnailkept grabbing while a human heldsuppressed: human_has_control; hero captions "Hidden while someone has control"request_handoff/takeover from the gateway or CLI process never reached the Desktop (on_changeis in-process)hermes servewatches each served profile'slease.json(0.5 s mtime poll, epoch-deduped) and broadcastsdisplay.leaseviewer_idclient-chosen and disclosed viadisplay.status: any client could co-drive orrelease(None)the holderdisplay.observereturns it; reuse only on the minting connection), snapshots/events carryviewer_hashnever the id; bare release refused unlessforcebot-desktopticket logged in on/api/ws_ws_auth_reasonrejects that providerxorg-x11-server-utils/xorg-x11-utils→ dnf5 aborts the whole installBINARY_PACKAGESmap + invariant test that every required binary maps into every distro list; apt gainsx11-xkb-utils, pacmandbus xorg-xpropclose(1000)on unmount only (@whyyagswhy); 1005/1006 keep the exclusion (documented contract test)--helpisEventForBotScreeneverywhere;?? 'local'; hero keyed by bot (@whyyagswhy):N; no per-profile start lock; pid recycled → wrong killpg;exec dbus-run-sessiondropped the Xvnc EXIT trap (orphan X server); log never rotatedstart.lock; pid + create_time; launcher stays supervisor; truncate per startshell=Truekillpg;claim()before the thread; CLI routed throughinstall_packageswait_for_humanran the full timeout when nobody answeredno_takeoveraftergrace(default 60 s)[]/null/unknown holder read as agent or raisedcontrol-takenoverlay unreachable; ⌘W on the canvas closed a terminal tab; sidebar row remounted every paint; older backends spun forever; olderdisplay.statusrolled back a newer lease; no way to reclaim after a reloadretainProfileacross install/attach; close code 4000;data-remote-screenswallows ⌘W; stable render identity;-32601→ settled; epoch ordering; Hand back (force)--remote-debugging-port=0; agent attaches via--cdpto the live instance (both launch orders live-proven)-SendCutText=0on XvncRulings (not changed, by design): no lease TTL/heartbeat (a vanished viewer is recovered by an explicit Hand back (force), never by the agent resuming on its own); takeover does not wait for in-flight input; SetEncodings needs no cap (16-bit count); Nous Cloud gets this with the next server release.
Live, integrated head, real Xvnc +
hermes serve+ headless Hermes Desktop over CDP: client-chosen viewer id ignored and a minted one returned; acquire →viewer_id: null+ matchingviewer_hash;display.statusdiscloses no id; thumbnail suppressed while human holds; bare release refused (viewer_mismatch), forced release works; display ticket on/api/ws→ HTTP 403; arequest_handofffrom a separate process arrived asdisplay.leasein 0.24 s; in the Desktop: Take over / Hand back round-trip, reload hands back (1000), a takeover made by another process repainted the pane with its reason and offered Hand back (force), which returned control.Human holds: red ring, "You are in control", the hero now says "Hidden while someone has control" instead of streaming their screen.
Another window took over (reason shown); this window offers "Hand back (force)".
Round 6 (independent review + design change)
request_handoff/wait_for_humanare gone fromcomputer_use, along with the lease'spending_handoff, the "Bot needs you" badge and the per-host schema rewriter. They blocked a tool call waiting for a human who, off the Desktop pane, was never watching, and fought the 420 s tool deadline. Takeover is human-initiated only.browser_vault_fill/enter_code/save_loginreach the page over the supervisor socket and skipped the fence; they now run under the samerun_fencedadmission + epoch check.--clone-allstrips the source's screen identity (P1):launcher.pid,env,rfb.sock, lease. A clone no longer believes it owns the source's X server (screen stopon the clone killed the source's desktop).setScreenLeaserecords a newer epoch even when nothing visible changed.display.stop/display.lease.releaseunder a human is decided inside the lease transition (a racing takeover can no longer be acknowledged and silently revoked).hermes profile renamewithout reseeding the panel layout.Not in this PR
The hosted/cloud-browser leg of #92524 and its dashboard integration; an agent-initiated "please take over" signal (removed in round 6: it only made sense with a person watching the Desktop pane; the bot asks in its reply instead); per-bot OS users (currently per-profile displays + XDG dirs under one user), Wayland compositors, macOS/Windows hosts (they have one real seat;
display.statusreportssupported: falseand the pane says so).Screenshots (Hermes Desktop, live against
hermes serveon this Linux host)Getting to a bot's computer
Bots → Hermes → Scheduled Jobs: the bot's screen is the hero at the very top of the pane, above the title and the routines (screen off, chip offers Start)
Screen running: the hero is a live picture of the bot's desktop (Nous wallpaper, dark theme, Chrome + terminal opened from the dock), refreshed every few seconds; caption Live · bot in control, chip Open live
Clicking the picture expands into the live Screen pane (noVNC canvas, Take over)
Sessions sidebar grouped by profile: a compact Screen row under the
defaultprofile header, so the profile's computer is reachable from its conversationsThe bot's desktop itself (what display.thumbnail returns): Nous gradient wallpaper, dark top bar with task list + clock, bottom dock of only the programs present on the host (Terminal, Browser pinned to the bot's Chrome); anything the user installs shows in the Applications menu
Installing the screen packages from the Desktop side
Host without TigerVNC/Xfce: the pane becomes an install card with the exact package command and Install on host
Install on host → the gateway host asks for its sudo password through the existing masked Administrator password card (sent to that host only, redacted from the trace)
Cancel → Install cancelled: no sudo password was provided; nothing was spawned, the card returns
Take over and hand back
Watching: Bot is in control, no border, keyboard and mouse ignored by the bridge
Take over: red border, chip You are in control, human typing into the bot's terminal on the Xfce desktop
The typed command executed on the bot's screen (echo HUMAN-TYPED-VIA-HERMES-DESKTOP); the chat composer stayed untouched
A second viewer took the lease: this pane shows Another viewer is in control and its RFB socket was closed with 4000 control-taken
Hand back: chip returns to Bot is in control, lease → agent, computer_use captures work again
Infographic