Repository navigation
feat: run real Cowork agent tasks on the demo box through an unprivileged host launcher - #870
Conversation
control-plane cannot exec Apptainer from its own container (musl base, no /dev/fuse, no CAP_SYS_ADMIN), and granting it those privileges would put sandbox-escape-class capability on the one process holding the payment and database secrets. Split the launcher out instead: agent-engine gains a -serve mode that runs the existing SandboxEngine as a long-lived host process behind a Unix socket, and control-plane reaches it through that socket when HIVE_AGENT_ENGINE_SOCKET is set. Also fixes the two things that made a real launch impossible even with every path variable set: a sandbox launched with --containall has no persisted agent profile store, so conversations now start with inline agent settings carrying the model endpoint, and the model host is added to the sandbox egress allowlist so the agent can reach it at all. Refs #780, #781
…killing a launch Two failures found on the first real launch against the demo box. The readiness check treated a dialable control socket as a ready server. The socat shim inside the image creates that socket immediately while the Python agent-server behind it takes tens of seconds to bind, so the first request after the wait died with a bare EOF. Readiness now requires a real HTTP response, and the default wait grew from 30 seconds to 3 minutes to cover a measured cold start. Task creation also ran the launch on the caller's request context, so closing the browser tab cancelled a live sandbox launch and the follow-up state write with it. The launch and both transitions now run on a detached context with their own timeout.
…r it works The notice was derived from any task in the list, so a task blocked before the runtime existed kept the warning on screen forever. It now reads the newest task only, which is the one that describes the current deployment.
Removes the temporary bring-up workflow now that the deploy job carries the same steps.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
|
Warning Review limit reached
Next review available in: 43 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (7)
📝 WalkthroughWalkthroughThe PR introduces a host-run agent-engine daemon with authenticated Unix-socket APIs. The control plane can use the daemon for sandbox tasks, while launch readiness, inline LLM settings, timeouts, deployment automation, and task-console warning behavior are updated. ChangesAgent-engine daemon and remote execution
Task console warning state
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant ControlPlane
participant RemoteAgentEngine
participant SandboxEngine
ControlPlane->>RemoteAgentEngine: Submit task over Unix socket
RemoteAgentEngine->>SandboxEngine: Launch sandbox session
SandboxEngine-->>RemoteAgentEngine: Return session reference
RemoteAgentEngine-->>ControlPlane: Return launch response
ControlPlane->>RemoteAgentEngine: Poll status or cancel session
RemoteAgentEngine-->>ControlPlane: Return session state
Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Same trust shape the internal RAG ingest and routing endpoints already carry: the value crosses a process boundary because a request context cannot, and the only caller fills it from the authenticated task row.
Resolves the .wolf/buglog.jsonl divergence that GitHub's server-side merge could not auto-resolve. The union merge driver is a local git setting, so the pull request showed as conflicting on GitHub and no pull_request workflow run was ever created for it.
There was a problem hiding this comment.
Actionable comments posted: 7
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@apps/agent-engine/cmd/agent-engine/serve.go`:
- Around line 96-107: Update the -serve configuration validation to require
HIVE_AGENT_ENGINE_LLM_MODEL unconditionally, removing profile-only acceptance.
In the serve setup flow around profileIDRaw and uuid.Parse, do not parse or pass
HIVE_AGENT_ENGINE_PROFILE_ID; allow a valid inline LLM configuration even when
the profile ID is stale or invalid.
- Around line 129-137: Update the resolveEgressHosts closure to copy the hosts
returned by egress.Effective before appending llmHost, ensuring the original
slice and its backing array are never mutated while preserving the existing
empty-llmHost return behavior.
- Around line 236-244: Update the http.Server initialization in the serve flow
to set bounded ReadTimeout and IdleTimeout values, limiting request-body reads
and idle connections. Keep WriteTimeout unset so the /launch handler can wait
for sandbox startup, and preserve the existing ReadHeaderTimeout and
MaxBytesReader behavior.
In `@apps/control-plane/internal/agentengine/remote.go`:
- Around line 138-146: Update the non-OK response handling in post to stop
returning the daemon’s decoded error text or raw response body. Return only a
stable local error containing the operation path and HTTP status, while
retaining any detailed daemon error exclusively through protected logging if
already available.
In `@deploy/docker/docker-compose.yml`:
- Around line 378-396: Update the standalone-service comment around the earlier
control-plane configuration to remove the claim that real tasks use an
in-process control-plane engine. State that real per-task launches use the host
daemon when HIVE_AGENT_ENGINE_SOCKET is set, consistent with the
HIVE_AGENT_ENGINE_SOCKET configuration and its forwarding behavior.
In `@docs/proof/agent-engine-live-2026-08-11/README.md`:
- Around line 25-29: Update the fenced timeline block in the README by adding
the text language identifier to its opening fence, while leaving the timeline
contents unchanged.
In `@scripts/install-agent-engine-host.sh`:
- Around line 101-117: Update the ENV_FILE generation block in
install-agent-engine-host.sh to shell-escape every emitted value before the file
is later sourced, including URLs, API keys, paths, tokens, and quota defaults.
Use Bash-safe serialization such as printf with %q (or replace sourcing with a
non-shell format) while preserving the existing variable names and values.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: b2a6ab8a-5c28-41dc-9e5d-c063afa658ad
⛔ Files ignored due to path filters (1)
docs/proof/agent-engine-live-2026-08-11/cowork-task-done.pngis excluded by!**/*.png
📒 Files selected for processing (18)
.env.example.github/workflows/deploy-demo-box.yml.wolf/buglog.jsonlapps/agent-console/components/task-console.test.tsxapps/agent-console/components/task-console.tsxapps/agent-engine/cmd/agent-engine/main.goapps/agent-engine/cmd/agent-engine/serve.goapps/agent-engine/internal/controlclient/client.goapps/agent-engine/internal/engine/engine.goapps/agent-engine/internal/engine/engine_test.goapps/control-plane/cmd/server/main.goapps/control-plane/internal/agentengine/remote.goapps/control-plane/internal/agenttask/service.godeploy/apptainer/README.mddeploy/docker/docker-compose.ymldocs/proof/agent-engine-live-2026-08-11/README.mdscripts/install-agent-engine-host.shtools/lint-no-direct-tenant-id.mjs
Seven review findings on the -serve daemon and its installer. The daemon no longer accepts an agent profile ID in place of a model alias. The sandbox runs with --containall, so the agent-server resolves a profile against a filesystem that is a fresh empty container every session and can never hold one, which made a profile-only launch report healthy and then fail every task. HIVE_AGENT_ENGINE_LLM_MODEL is now required outright. The egress resolver copies the allowlist before appending the model host, rather than writing into the backing array egress.Effective returned. The daemon's http.Server gains bounded ReadTimeout and IdleTimeout. WriteTimeout stays unset because /launch legitimately holds its response open while a cold sandbox starts. The control-plane client no longer returns the daemon's error text, which names sandbox paths, egress hosts and the model provider behind a failed call. That detail is logged server-side and callers get a stable error carrying only the operation and the status code, matching the provider-blind rule both service boundaries already follow. A regression test covers the boundary. The installer serializes engine.env through printf %q. The unit entry point sources that file, so a value containing whitespace, a quote, a newline or a command substitution would otherwise execute as the deploying user. Two documentation fixes: the compose comment claimed real tasks use the in-process engine, which the socket-based daemon replaced, and the proof timeline block had no language for markdownlint.
|
All seven review findings are addressed in 92ed9f9.
Separately, this pull request had never received a real |
…ature branches (#874) Documentation and protocol only. No runtime behaviour changes. The two `.wolf/hooks` edits are message strings. ## 1. CLAUDE.md and DEMO.md named the wrong agent-engine gate `deploy/apptainer/README.md` was corrected by #870, but `CLAUDE.md` and `DEMO.md` still said `HIVE_AGENT_SIF_PATH` is what makes an agent task launch. They contradicted both the README and the code. This is worse than an ordinary stale document. `CLAUDE.md` is loaded into every agent's context in every session, so a wrong statement in it propagates into new briefs indefinitely. That exact mechanism kept a revoked business rule alive for nine weeks in this repository. Re-derived from `buildAgentEngine` in `apps/control-plane/cmd/server/main.go` at `c1ca04d5` rather than from the previous documentation: * **Socket arm, checked first, and what the demo box actually runs.** When `HIVE_AGENT_ENGINE_SOCKET` is set, control-plane hands every launch to the unprivileged host launcher over that Unix socket, authenticating with `CONTROL_PLANE_INTERNAL_TOKEN`, and none of the path variables are read at all. Both files now lead with this. * **In-process arm.** Needs a non-nil egress service, so a live DB pool, plus five variables, all of them: `HIVE_AGENT_ENGINE_SIF_PATH`, `HIVE_AGENT_ENGINE_PACKS_DIR`, `HIVE_AGENT_ENGINE_WORKSPACE_ROOT`, `HIVE_AGENT_ENGINE_RUN_DIR`, `HIVE_AGENT_ENGINE_PROFILE_ID`. Missing any one falls back to `NotConfiguredEngine`, every submitted task fails immediately, and the boot WARN names what was missing. Under docker compose this arm cannot succeed regardless of the variables. * **Defaulted, so optional:** `HIVE_AGENT_ENGINE_SESSION_API_KEY`, `HIVE_QUOTA_TENANT_CONCURRENCY` (4), `HIVE_QUOTA_USER_CONCURRENCY` (2), `HIVE_SANDBOX_MEMORY_LIMIT` (4G), `HIVE_SANDBOX_CPU_LIMIT` (2), `HIVE_SANDBOX_PIDS_LIMIT` (512). * **The daemon reads its own set** in `apps/agent-engine/cmd/agent-engine/serve.go`, including the three `HIVE_AGENT_ENGINE_LLM_*` variables. * **`HIVE_AGENT_SIF_PATH` gates nothing.** It survives only as the `-sif` flag default at `apps/agent-engine/cmd/agent-engine/main.go:41` and for the compose smoke-test service under `--profile agent`. The stale `.env.example` block that repeated the same claim ("refuses to start without it") is corrected in the same way. The socket variables were already documented higher up that file by #870. ## 2. Buglog entries move off feature branches, adopting #873 Appending to the tracked `.wolf/buglog.jsonl` from a feature branch has two proven failure modes: every parallel fix pull request conflicts serially with every other one, and an unmergeable branch produces **zero** `pull_request` CI runs, which reads as a broken workflow rather than an unmergeable branch. Both come from concurrent writes to one file. `merge=union` stays in `.gitattributes` because it still resolves concurrent appends in local merges and rebases. The protocol now states plainly that GitHub's server-side merge ignores it, so nobody re-derives the trap. **Recording a bug entry is still mandatory.** Only the destination changed: 1. While the fix is in flight, the entry rides in the fix pull request body under a "Buglog entry" heading, as the JSON line to be appended. It is reviewable there and attached to the fix. 2. After the fix merges, it lands on `main` through a separate pull request whose diff is `.wolf/buglog.jsonl` and nothing else. One such pull request open at a time across all agents, and batching several entries into one is preferred. That route is manual, and the protocol says so rather than describing a process nobody can perform. The post-merge automation that would replace it does not exist yet and is tracked in #873. Those pull requests merge cheaply because `.wolf/*` is on the inert-path allowlist in `.github/workflows/ci.yml`, so the six required checks report green without running their heavy steps. The rule also warns that `bugstore.js add` writes the tracked file directly, so running it on a fix branch dirties the working tree and the next `git add -A` sweeps it into the fix commit. Updated wherever the protocol is stated: `.claude/rules/openwolf.md`, the OpenWolf section of `CLAUDE.md`, `.claude/skills/memory-tools.md` (which also still pointed at the gitignored `buglog.json`), and the two `.wolf/hooks` nudges that previously told an agent to append the line right where it was standing. ## Verification * Every claim above read out of the code at `c1ca04d5`, not from the documentation being replaced. * `.wolf/*` allowlist and the `if: always()` required jobs confirmed in `.github/workflows/ci.yml`, which is what makes the buglog-only pull request route cheap. * Both edited hooks pass `node --check`. No test asserts either message string. * `npm run lint:proof-tokens` passes. No credential appears in this diff; variables are referenced by name only. * No UI surface is touched, so no screenshot applies. Note that this pull request edits two `.wolf/hooks/*.js` files, which the CI path filter denies by executable extension, so it deliberately runs the full required suite rather than taking the inert-path skip. Closes #873 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Clarified agent-engine setup, host launcher configuration, socket-based task execution, and standalone smoke-test behavior. * Updated environment-variable guidance, including the limited role of the agent-engine SIF path. * Documented bug-log workflows, merge-conflict handling, pull-request coordination, and feature-branch requirements. * Revised automated reminders to record feature-branch bug fixes in pull-request details. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… task lifecycle Rewrites the interaction-coverage probe so it can actually fail. Authentication now comes from tests/e2e/support/live-auth.ts, which mints a session through the admin one-time-token flow. It needs no password and changes none. The previous version skipped fourteen of twenty-two controls on the grounds that authenticating required rotating the shared demo account's password through an admin endpoint that was returning 500s. Rotation is forbidden, but it was never the only way in, and that reason was baked into a committed coverage file. Rewrites the task lifecycle assertions against what the deployment does since PR #870. A task now progresses queued, running, done through an unprivileged host launcher rather than resolving straight to a terminal blocked state, so the engine-unavailable notice must be absent and the cancel button must be offered. C17 asserts the button cancels a live task and that a second cancel on the now-terminal task is refused with exactly 409, instead of the previous range assertion that also passed on a 404 or a 401. C18 no longer disables itself after its first run. It stubs an empty list rather than requiring the live account to have no history, which earlier controls in the same file guarantee it does. C23 is new: it enumerates every focusable control the deployed sign-in and tasks screens render and requires the set to equal the ledger's dom entries, so a control cannot ship without raising the denominator. It also asserts the absence of the transcript pane the ledger claims does not exist yet. The coverage builder reads Playwright's own per-test verdict rather than the first attempt of a retried test, and treats a retry-pass as unproven. It exits non-zero when a control has no test at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… task lifecycle Rewrites the interaction-coverage probe so it can actually fail. Authentication now comes from tests/e2e/support/live-auth.ts, which mints a session through the admin one-time-token flow. It needs no password and changes none. The previous version skipped fourteen of twenty-two controls on the grounds that authenticating required rotating the shared demo account's password through an admin endpoint that was returning 500s. Rotation is forbidden, but it was never the only way in, and that reason was baked into a committed coverage file. Rewrites the task lifecycle assertions against what the deployment does since PR #870. A task now progresses queued, running, done through an unprivileged host launcher rather than resolving straight to a terminal blocked state, so the engine-unavailable notice must be absent and the cancel button must be offered. C17 asserts the button cancels a live task and that a second cancel on the now-terminal task is refused with exactly 409, instead of the previous range assertion that also passed on a 404 or a 401. C18 no longer disables itself after its first run. It stubs an empty list rather than requiring the live account to have no history, which earlier controls in the same file guarantee it does. C23 is new: it enumerates every focusable control the deployed sign-in and tasks screens render and requires the set to equal the ledger's dom entries, so a control cannot ship without raising the denominator. It also asserts the absence of the transcript pane the ledger claims does not exist yet. The coverage builder reads Playwright's own per-test verdict rather than the first attempt of a retried test, and treats a retry-pass as unproven. It exits non-zero when a control has no test at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… task lifecycle Rewrites the interaction-coverage probe so it can actually fail. Authentication now comes from tests/e2e/support/live-auth.ts, which mints a session through the admin one-time-token flow. It needs no password and changes none. The previous version skipped fourteen of twenty-two controls on the grounds that authenticating required rotating the shared demo account's password through an admin endpoint that was returning 500s. Rotation is forbidden, but it was never the only way in, and that reason was baked into a committed coverage file. Rewrites the task lifecycle assertions against what the deployment does since PR #870. A task now progresses queued, running, done through an unprivileged host launcher rather than resolving straight to a terminal blocked state, so the engine-unavailable notice must be absent and the cancel button must be offered. C17 asserts the button cancels a live task and that a second cancel on the now-terminal task is refused with exactly 409, instead of the previous range assertion that also passed on a 404 or a 401. C18 no longer disables itself after its first run. It stubs an empty list rather than requiring the live account to have no history, which earlier controls in the same file guarantee it does. C23 is new: it enumerates every focusable control the deployed sign-in and tasks screens render and requires the set to equal the ledger's dom entries, so a control cannot ship without raising the denominator. It also asserts the absence of the transcript pane the ledger claims does not exist yet. The coverage builder reads Playwright's own per-test verdict rather than the first attempt of a retried test, and treats a retry-pass as unproven. It exits non-zero when a control has no test at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Cowork tasks now actually run on the demo box. Before this branch every task on that deployment ended
Blockedwith "agent engine is not available on this deployment"; a task submitted from the UI now reachesDonewith output the agent produced inside its own Apptainer sandbox.Closes #780. Closes #781.
Proof
docs/proof/agent-engine-live-2026-08-11/cowork-task-done.png(see the README beside it), captured athttps://chat-hive.scubed.co/agent-workspaceafter deploying this branch to the box:Observed timeline for that run:
Deliberately verified by behaviour rather than by a green deploy log, per #869. Three things in that screenshot cannot occur on the previous build: a task in
Doneat all, a result string that only a live sandbox plus a live model call can produce, and the absence of the "runtime is not configured" banner that the old console derived from any blocked task in the list.The substrate decision
control-planeruns in an Alpine container. It cannot exec the host's Apptainer: no glibc loader for that binary, no/dev/fuse, and noCAP_SYS_ADMIN-class privilege. Granting that container those privileges was considered and refused. It is the one process holding the Stripe keys, the Supabase service-role key and the platform database DSN, so sandbox-escape-class privilege there is the worst possible placement on the box.The launcher runs on the host instead, as an ordinary unprivileged user, and
control-planereaches it over a Unix socket bind-mounted into the container. No capability is added to any container, nothing runs as root, andcontrol-plane's own capability set is untouched. The socket is created0600inside a0700directory, and the existing shared internal token is checked on every call so that bind-mounting it somewhere else later cannot silently widen the boundary.The box already had Apptainer 1.5.3 installed,
/dev/fusepresent, and unprivileged user namespaces enabled, so nothing needed installing as root. What was missing was the image and a process allowed to launch it.What was actually wrong, in the order it was found
HIVE_AGENT_ENGINE_*configuration on the box at all (Apptainer docs tell you to set HIVE_AGENT_SIF_PATH, but real task launches read HIVE_AGENT_ENGINE_SIF_PATH #781), and.env.exampledescribed a variable set that was both wrong and insufficient. Rewritten from the code: the socket variable is the only onecontrol-planeneeds on a containerised deployment, and the launcher reads the paths, the model settings, the quota ceilings and the sandbox limits.agent-enginegains a-servemode that runs the existingSandboxEnginebehind a Unix socket;control-planegains the client half. The engine, the quota manager, the egress proxy and the launcher are unchanged and still run exactly once.agent_profile_id, but a sandbox launched with--containallstarts from an empty container filesystem every session, so that lookup could only ever returnProfileNotFound. Conversations now start with inline agent settings carrying the model endpoint, which is the only shape that resolves there.WaitReadyaccepted a dialable control socket, but the socat shim inside the image creates that socket immediately while the Python agent-server behind it takes tens of seconds to bind. The first request landed in that window and died with a bareEOF. Readiness now requires a real HTTP response, and the default wait went from 30 seconds to 3 minutes to cover a measured cold start.CreateTaskran the launch on the caller's request context, so navigating away cancelled a running sandbox launch, and the follow-up state write with it. Both now run on a detached context with their own timeout.Deploy
scripts/install-agent-engine-host.shis idempotent and runs on every deploy: it fetches the CI-built.sifwhen the host has none, builds the launcher binary in the same Go image the rest of the stack uses, writes a0600env file, and restarts the systemd user unit. The deploy job exports the socket variables for compose interpolation and appends them to the box's.envwhen absent, so a compose run by hand on the box behaves like the deploy rather than quietly coming up unconfigured.No secret value appears in any workflow file, script, log line or committed file. The model key is
secrets.HIVE_API_KEY, referenced by name and written only into the daemon's own0600env file.Tests
engine: inline agent settings are sent when a model is configured, and the profile path still applies when one is not.agent-console: an older blocked task no longer contradicts a newer task that ran.agent-engineandagenttasksuites pass, plus the 26 existing agent-console console tests.Follow-ups, not in scope here
#469(whether the orchestration loop should live outside the sandbox) is untouched. The loop still runs inside.🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Bug Fixes
Documentation