Skip to content

feat: run real Cowork agent tasks on the demo box through an unprivileged host launcher - #870

Merged
sakibsadmanshajib merged 16 commits into
mainfrom
feat/agent-engine-demo-box
Aug 11, 2026
Merged

sakibsadmanshajib merged 16 commits into
mainfrom
feat/agent-engine-demo-box

Conversation

@sakibsadmanshajib

@sakibsadmanshajib sakibsadmanshajib commented Aug 11, 2026 •

Copy link
Copy Markdown
Owner

Cowork tasks now actually run on the demo box. Before this branch every task on that deployment ended Blocked with "agent engine is not available on this deployment"; a task submitted from the UI now reaches Done with output the agent produced inside its own Apptainer sandbox.

Closes #780. Closes #781.

Proof

docs/proof/agent-engine-live-2026-08-11/cowork-task-done.png (see the README beside it), captured at https://chat-hive.scubed.co/agent-workspace after deploying this branch to the box:

Cowork task Done with a real result

Observed timeline for that run:

12:04:33  Queued    waiting for a sandbox
12:04:56  Running   sandbox up, conversation started
12:20:29  Done      result returned

Deliberately verified by behaviour rather than by a green deploy log, per #869. Three things in that screenshot cannot occur on the previous build: a task in Done at all, a result string that only a live sandbox plus a live model call can produce, and the absence of the "runtime is not configured" banner that the old console derived from any blocked task in the list.

The substrate decision

control-plane runs in an Alpine container. It cannot exec the host's Apptainer: no glibc loader for that binary, no /dev/fuse, and no CAP_SYS_ADMIN-class privilege. Granting that container those privileges was considered and refused. It is the one process holding the Stripe keys, the Supabase service-role key and the platform database DSN, so sandbox-escape-class privilege there is the worst possible placement on the box.

The launcher runs on the host instead, as an ordinary unprivileged user, and control-plane reaches it over a Unix socket bind-mounted into the container. No capability is added to any container, nothing runs as root, and control-plane's own capability set is untouched. The socket is created 0600 inside a 0700 directory, and the existing shared internal token is checked on every call so that bind-mounting it somewhere else later cannot silently widen the boundary.

The box already had Apptainer 1.5.3 installed, /dev/fuse present, and unprivileged user namespaces enabled, so nothing needed installing as root. What was missing was the image and a process allowed to launch it.

What was actually wrong, in the order it was found

  1. No HIVE_AGENT_ENGINE_* configuration on the box at all (Apptainer docs tell you to set HIVE_AGENT_SIF_PATH, but real task launches read HIVE_AGENT_ENGINE_SIF_PATH #781), and .env.example described a variable set that was both wrong and insufficient. Rewritten from the code: the socket variable is the only one control-plane needs on a containerised deployment, and the launcher reads the paths, the model settings, the quota ceilings and the sandbox limits.
  2. No substrate (control-plane's compose container cannot exec a real Apptainer sandbox even with all HIVE_AGENT_ENGINE_* vars set #780). Fixed as above. agent-engine gains a -serve mode that runs the existing SandboxEngine behind a Unix socket; control-plane gains the client half. The engine, the quota manager, the egress proxy and the launcher are unchanged and still run exactly once.
  3. No agent profile can exist inside the sandbox. Launches referenced a stored agent_profile_id, but a sandbox launched with --containall starts from an empty container filesystem every session, so that lookup could only ever return ProfileNotFound. Conversations now start with inline agent settings carrying the model endpoint, which is the only shape that resolves there.
  4. The model endpoint was not reachable. Sandbox egress is deny-all by default for a tenant with no policy row, and the model call goes through the same proxy. The launcher appends the model host, and only that host, to whatever the tenant policy resolves to. It is Hive's own metered gateway, not a tenant-chosen destination, and without it no task can run at all.
  5. Rootless cgroups need a systemd user session. Apptainer enforces the per-session memory, CPU and PID limits through cgroups, which fails with "cannot use cgroups, DBUS_SESSION_BUS_ADDRESS is not set" when launched from a bare CI step. The daemon runs as a transient systemd user unit, which supplies the delegation and also outlives the job that starts it.
  6. Readiness was measured on the wrong thing. WaitReady accepted a dialable control socket, but the socat shim inside the image creates that socket immediately while the Python agent-server behind it takes tens of seconds to bind. The first request landed in that window and died with a bare EOF. Readiness now requires a real HTTP response, and the default wait went from 30 seconds to 3 minutes to cover a measured cold start.
  7. A closed browser tab killed live launches. CreateTask ran the launch on the caller's request context, so navigating away cancelled a running sandbox launch, and the follow-up state write with it. Both now run on a detached context with their own timeout.

Deploy

scripts/install-agent-engine-host.sh is idempotent and runs on every deploy: it fetches the CI-built .sif when the host has none, builds the launcher binary in the same Go image the rest of the stack uses, writes a 0600 env file, and restarts the systemd user unit. The deploy job exports the socket variables for compose interpolation and appends them to the box's .env when absent, so a compose run by hand on the box behaves like the deploy rather than quietly coming up unconfigured.

No secret value appears in any workflow file, script, log line or committed file. The model key is secrets.HIVE_API_KEY, referenced by name and written only into the daemon's own 0600 env file.

Tests

  • engine: inline agent settings are sent when a model is configured, and the profile path still applies when one is not.
  • agent-console: an older blocked task no longer contradicts a newer task that ran.
  • Full agent-engine and agenttask suites pass, plus the 26 existing agent-console console tests.

Follow-ups, not in scope here

  • #469 (whether the orchestration loop should live outside the sandbox) is untouched. The loop still runs inside.
  • The launcher holds its session registry in memory, so a daemon restart loses the ability to poll or cancel an in-flight session. The sandbox is a child process and dies with it either way, so this only matters once sandboxes outlive the launcher.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added a host-based agent engine daemon for launching, monitoring, cancelling, and health-checking sandbox sessions.
    • Added Unix-socket communication between the control plane and agent engine, with optional inline model configuration.
    • Added an automated host installation and service setup workflow.
  • Bug Fixes

    • Improved readiness checks and increased startup timeouts.
    • Prevented task launches from remaining stuck after client disconnects.
    • Fixed outdated engine-configuration warnings in the task console.
  • Documentation

    • Updated deployment, environment configuration, and live-demo documentation.

control-plane cannot exec Apptainer from its own container (musl base, no
/dev/fuse, no CAP_SYS_ADMIN), and granting it those privileges would put
sandbox-escape-class capability on the one process holding the payment and
database secrets. Split the launcher out instead: agent-engine gains a -serve
mode that runs the existing SandboxEngine as a long-lived host process behind
a Unix socket, and control-plane reaches it through that socket when
HIVE_AGENT_ENGINE_SOCKET is set.

Also fixes the two things that made a real launch impossible even with every
path variable set: a sandbox launched with --containall has no persisted
agent profile store, so conversations now start with inline agent settings
carrying the model endpoint, and the model host is added to the sandbox
egress allowlist so the agent can reach it at all.

Refs #780, #781
…killing a launch

Two failures found on the first real launch against the demo box.

The readiness check treated a dialable control socket as a ready server. The
socat shim inside the image creates that socket immediately while the Python
agent-server behind it takes tens of seconds to bind, so the first request
after the wait died with a bare EOF. Readiness now requires a real HTTP
response, and the default wait grew from 30 seconds to 3 minutes to cover a
measured cold start.

Task creation also ran the launch on the caller's request context, so closing
the browser tab cancelled a live sandbox launch and the follow-up state write
with it. The launch and both transitions now run on a detached context with
their own timeout.
…r it works

The notice was derived from any task in the list, so a task blocked before
the runtime existed kept the warning on screen forever. It now reads the
newest task only, which is the one that describes the current deployment.
Removes the temporary bring-up workflow now that the deploy job carries the
same steps.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@cursor

cursor Bot commented Aug 11, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@coderabbitai

coderabbitai Bot commented Aug 11, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@sakibsadmanshajib, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 43 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 41bf9cb9-1790-4fd3-b685-40b0e821b97d

📥 Commits

Reviewing files that changed from the base of the PR and between 27f97ce and 92ed9f9.

📒 Files selected for processing (7)
  • .wolf/buglog.jsonl
  • apps/agent-engine/cmd/agent-engine/serve.go
  • apps/control-plane/internal/agentengine/remote.go
  • apps/control-plane/internal/agentengine/remote_test.go
  • deploy/docker/docker-compose.yml
  • docs/proof/agent-engine-live-2026-08-11/README.md
  • scripts/install-agent-engine-host.sh
📝 Walkthrough

Walkthrough

The PR introduces a host-run agent-engine daemon with authenticated Unix-socket APIs. The control plane can use the daemon for sandbox tasks, while launch readiness, inline LLM settings, timeouts, deployment automation, and task-console warning behavior are updated.

Changes

Agent-engine daemon and remote execution

Layer / File(s) Summary
Daemon runtime and launch contracts
apps/agent-engine/cmd/agent-engine/*, apps/agent-engine/internal/controlclient/client.go, apps/agent-engine/internal/engine/*
agent-engine -serve exposes authenticated launch, status, cancellation, and health endpoints over a Unix socket. Readiness uses HTTP health responses. Launches support inline LLM settings or stored agent profiles.
Remote control-plane execution and task lifecycle
apps/control-plane/internal/agentengine/remote.go, apps/control-plane/cmd/server/main.go, apps/control-plane/internal/agenttask/service.go
The control plane selects a remote agent engine when configured. Remote requests use authenticated JSON over the Unix socket. Launches and task transitions use independent five-minute timeouts.
Host installation and deployment wiring
scripts/install-agent-engine-host.sh, .github/workflows/deploy-demo-box.yml, .env.example, deploy/docker/docker-compose.yml, deploy/apptainer/README.md, tools/lint-no-direct-tenant-id.mjs, .wolf/buglog.jsonl, docs/proof/agent-engine-live-2026-08-11/README.md
The host installer builds, configures, starts, and verifies the daemon. Deployment exports socket settings and credentials. Compose, environment examples, documentation, lint configuration, and deployment records describe the new topology.

Task console warning state

Layer / File(s) Summary
Newest-task warning behavior
apps/agent-console/components/task-console.tsx, apps/agent-console/components/task-console.test.tsx
The engine-configuration warning now reflects only the newest task. A regression test covers newer successful tasks after older engine-unavailable tasks.

Estimated code review effort: 4 (Complex) | ~60 minutes

Sequence Diagram(s)

sequenceDiagram
  participant ControlPlane
  participant RemoteAgentEngine
  participant SandboxEngine
  ControlPlane->>RemoteAgentEngine: Submit task over Unix socket
  RemoteAgentEngine->>SandboxEngine: Launch sandbox session
  SandboxEngine-->>RemoteAgentEngine: Return session reference
  RemoteAgentEngine-->>ControlPlane: Return launch response
  ControlPlane->>RemoteAgentEngine: Poll status or cancel session
  RemoteAgentEngine-->>ControlPlane: Return session state
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The changes address #780 with a host-side Unix-socket launcher and address #781 with corrected HIVE_AGENT_ENGINE_* documentation and deployment wiring.
Out of Scope Changes check ✅ Passed The reviewed changes support the linked issues through launcher implementation, readiness fixes, task handling, deployment wiring, tests, and documentation.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: running real Cowork agent tasks on the demo box through an unprivileged host launcher.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/agent-engine-demo-box

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Same trust shape the internal RAG ingest and routing endpoints already carry:
the value crosses a process boundary because a request context cannot, and the
only caller fills it from the authenticated task row.
Resolves the .wolf/buglog.jsonl divergence that GitHub's server-side merge
could not auto-resolve. The union merge driver is a local git setting, so the
pull request showed as conflicting on GitHub and no pull_request workflow run
was ever created for it.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/agent-engine/cmd/agent-engine/serve.go`:
- Around line 96-107: Update the -serve configuration validation to require
HIVE_AGENT_ENGINE_LLM_MODEL unconditionally, removing profile-only acceptance.
In the serve setup flow around profileIDRaw and uuid.Parse, do not parse or pass
HIVE_AGENT_ENGINE_PROFILE_ID; allow a valid inline LLM configuration even when
the profile ID is stale or invalid.
- Around line 129-137: Update the resolveEgressHosts closure to copy the hosts
returned by egress.Effective before appending llmHost, ensuring the original
slice and its backing array are never mutated while preserving the existing
empty-llmHost return behavior.
- Around line 236-244: Update the http.Server initialization in the serve flow
to set bounded ReadTimeout and IdleTimeout values, limiting request-body reads
and idle connections. Keep WriteTimeout unset so the /launch handler can wait
for sandbox startup, and preserve the existing ReadHeaderTimeout and
MaxBytesReader behavior.

In `@apps/control-plane/internal/agentengine/remote.go`:
- Around line 138-146: Update the non-OK response handling in post to stop
returning the daemon’s decoded error text or raw response body. Return only a
stable local error containing the operation path and HTTP status, while
retaining any detailed daemon error exclusively through protected logging if
already available.

In `@deploy/docker/docker-compose.yml`:
- Around line 378-396: Update the standalone-service comment around the earlier
control-plane configuration to remove the claim that real tasks use an
in-process control-plane engine. State that real per-task launches use the host
daemon when HIVE_AGENT_ENGINE_SOCKET is set, consistent with the
HIVE_AGENT_ENGINE_SOCKET configuration and its forwarding behavior.

In `@docs/proof/agent-engine-live-2026-08-11/README.md`:
- Around line 25-29: Update the fenced timeline block in the README by adding
the text language identifier to its opening fence, while leaving the timeline
contents unchanged.

In `@scripts/install-agent-engine-host.sh`:
- Around line 101-117: Update the ENV_FILE generation block in
install-agent-engine-host.sh to shell-escape every emitted value before the file
is later sourced, including URLs, API keys, paths, tokens, and quota defaults.
Use Bash-safe serialization such as printf with %q (or replace sourcing with a
non-shell format) while preserving the existing variable names and values.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: b2a6ab8a-5c28-41dc-9e5d-c063afa658ad

📥 Commits

Reviewing files that changed from the base of the PR and between a1cf17a and 27f97ce.

⛔ Files ignored due to path filters (1)
  • docs/proof/agent-engine-live-2026-08-11/cowork-task-done.png is excluded by !**/*.png
📒 Files selected for processing (18)
  • .env.example
  • .github/workflows/deploy-demo-box.yml
  • .wolf/buglog.jsonl
  • apps/agent-console/components/task-console.test.tsx
  • apps/agent-console/components/task-console.tsx
  • apps/agent-engine/cmd/agent-engine/main.go
  • apps/agent-engine/cmd/agent-engine/serve.go
  • apps/agent-engine/internal/controlclient/client.go
  • apps/agent-engine/internal/engine/engine.go
  • apps/agent-engine/internal/engine/engine_test.go
  • apps/control-plane/cmd/server/main.go
  • apps/control-plane/internal/agentengine/remote.go
  • apps/control-plane/internal/agenttask/service.go
  • deploy/apptainer/README.md
  • deploy/docker/docker-compose.yml
  • docs/proof/agent-engine-live-2026-08-11/README.md
  • scripts/install-agent-engine-host.sh
  • tools/lint-no-direct-tenant-id.mjs

Comment thread apps/agent-engine/cmd/agent-engine/serve.go Outdated
Comment thread apps/agent-engine/cmd/agent-engine/serve.go Outdated
Comment thread apps/agent-engine/cmd/agent-engine/serve.go Outdated
Comment thread apps/control-plane/internal/agentengine/remote.go Outdated
Comment thread deploy/docker/docker-compose.yml
Comment thread docs/proof/agent-engine-live-2026-08-11/README.md Outdated
Comment thread scripts/install-agent-engine-host.sh Outdated
Seven review findings on the -serve daemon and its installer.

The daemon no longer accepts an agent profile ID in place of a model
alias. The sandbox runs with --containall, so the agent-server resolves a
profile against a filesystem that is a fresh empty container every
session and can never hold one, which made a profile-only launch report
healthy and then fail every task. HIVE_AGENT_ENGINE_LLM_MODEL is now
required outright.

The egress resolver copies the allowlist before appending the model host,
rather than writing into the backing array egress.Effective returned.

The daemon's http.Server gains bounded ReadTimeout and IdleTimeout.
WriteTimeout stays unset because /launch legitimately holds its response
open while a cold sandbox starts.

The control-plane client no longer returns the daemon's error text, which
names sandbox paths, egress hosts and the model provider behind a failed
call. That detail is logged server-side and callers get a stable error
carrying only the operation and the status code, matching the
provider-blind rule both service boundaries already follow. A regression
test covers the boundary.

The installer serializes engine.env through printf %q. The unit entry
point sources that file, so a value containing whitespace, a quote, a
newline or a command substitution would otherwise execute as the
deploying user.

Two documentation fixes: the compose comment claimed real tasks use the
in-process engine, which the socket-based daemon replaced, and the proof
timeline block had no language for markdownlint.
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

All seven review findings are addressed in 92ed9f9.

  • serve.go daemon config: HIVE_AGENT_ENGINE_LLM_MODEL is now required outright and the profile ID is no longer parsed or passed in -serve mode. The sandbox runs with --containall, so the contained agent server can never resolve a profile, which is exactly the healthy-then-always-fails shape the review described.
  • serve.go egress hosts: the allowlist is copied before the model host is appended, so egress.Effective's backing array is never written through.
  • serve.go server timeouts: bounded ReadTimeout and IdleTimeout added. WriteTimeout stays unset, with a comment, because /launch legitimately holds its response open while a cold sandbox starts.
  • remote.go error boundary: the daemon's error text is logged server-side only. Callers now receive a stable error carrying just the operation and the status code, matching the provider-blind rule both boundaries already follow. A regression test covers it.
  • install-agent-engine-host.sh: every assignment in engine.env is serialized through printf %q, so a value containing whitespace, a quote, a newline, or a command substitution cannot execute when the unit entry point sources that file.
  • docker-compose.yml: the stale comment claiming real tasks use the in-process engine is corrected to describe the socket-based host daemon.
  • Proof README: the timeline fence is labelled text.

Separately, this pull request had never received a real pull_request CI run. The branch was unmergeable on a .wolf/buglog.jsonl divergence, and merge=union is a local git setting that GitHub's server-side merge does not apply, so refs/pull/870/merge could not be built and no run was ever created. Merging main in fixed that. The green now on record comes from a real pull_request run, and Live integration was opted in with the label rather than left skipped.

sakibsadmanshajib added a commit that referenced this pull request Aug 11, 2026
…ature branches (#874)

Documentation and protocol only. No runtime behaviour changes. The two
`.wolf/hooks` edits are message strings.

## 1. CLAUDE.md and DEMO.md named the wrong agent-engine gate

`deploy/apptainer/README.md` was corrected by #870, but `CLAUDE.md` and
`DEMO.md` still said `HIVE_AGENT_SIF_PATH` is what makes an agent task
launch. They contradicted both the README and the code.

This is worse than an ordinary stale document. `CLAUDE.md` is loaded
into every agent's context in every session, so a wrong statement in it
propagates into new briefs indefinitely. That exact mechanism kept a
revoked business rule alive for nine weeks in this repository.

Re-derived from `buildAgentEngine` in
`apps/control-plane/cmd/server/main.go` at `c1ca04d5` rather than from
the previous documentation:

* **Socket arm, checked first, and what the demo box actually runs.**
When `HIVE_AGENT_ENGINE_SOCKET` is set, control-plane hands every launch
to the unprivileged host launcher over that Unix socket, authenticating
with `CONTROL_PLANE_INTERNAL_TOKEN`, and none of the path variables are
read at all. Both files now lead with this.
* **In-process arm.** Needs a non-nil egress service, so a live DB pool,
plus five variables, all of them: `HIVE_AGENT_ENGINE_SIF_PATH`,
`HIVE_AGENT_ENGINE_PACKS_DIR`, `HIVE_AGENT_ENGINE_WORKSPACE_ROOT`,
`HIVE_AGENT_ENGINE_RUN_DIR`, `HIVE_AGENT_ENGINE_PROFILE_ID`. Missing any
one falls back to `NotConfiguredEngine`, every submitted task fails
immediately, and the boot WARN names what was missing. Under docker
compose this arm cannot succeed regardless of the variables.
* **Defaulted, so optional:** `HIVE_AGENT_ENGINE_SESSION_API_KEY`,
`HIVE_QUOTA_TENANT_CONCURRENCY` (4), `HIVE_QUOTA_USER_CONCURRENCY` (2),
`HIVE_SANDBOX_MEMORY_LIMIT` (4G), `HIVE_SANDBOX_CPU_LIMIT` (2),
`HIVE_SANDBOX_PIDS_LIMIT` (512).
* **The daemon reads its own set** in
`apps/agent-engine/cmd/agent-engine/serve.go`, including the three
`HIVE_AGENT_ENGINE_LLM_*` variables.
* **`HIVE_AGENT_SIF_PATH` gates nothing.** It survives only as the
`-sif` flag default at `apps/agent-engine/cmd/agent-engine/main.go:41`
and for the compose smoke-test service under `--profile agent`.

The stale `.env.example` block that repeated the same claim ("refuses to
start without it") is corrected in the same way. The socket variables
were already documented higher up that file by #870.

## 2. Buglog entries move off feature branches, adopting #873

Appending to the tracked `.wolf/buglog.jsonl` from a feature branch has
two proven failure modes: every parallel fix pull request conflicts
serially with every other one, and an unmergeable branch produces
**zero** `pull_request` CI runs, which reads as a broken workflow rather
than an unmergeable branch. Both come from concurrent writes to one
file.

`merge=union` stays in `.gitattributes` because it still resolves
concurrent appends in local merges and rebases. The protocol now states
plainly that GitHub's server-side merge ignores it, so nobody re-derives
the trap.

**Recording a bug entry is still mandatory.** Only the destination
changed:

1. While the fix is in flight, the entry rides in the fix pull request
body under a "Buglog entry" heading, as the JSON line to be appended. It
is reviewable there and attached to the fix.
2. After the fix merges, it lands on `main` through a separate pull
request whose diff is `.wolf/buglog.jsonl` and nothing else. One such
pull request open at a time across all agents, and batching several
entries into one is preferred.

That route is manual, and the protocol says so rather than describing a
process nobody can perform. The post-merge automation that would replace
it does not exist yet and is tracked in #873. Those pull requests merge
cheaply because `.wolf/*` is on the inert-path allowlist in
`.github/workflows/ci.yml`, so the six required checks report green
without running their heavy steps. The rule also warns that `bugstore.js
add` writes the tracked file directly, so running it on a fix branch
dirties the working tree and the next `git add -A` sweeps it into the
fix commit.

Updated wherever the protocol is stated: `.claude/rules/openwolf.md`,
the OpenWolf section of `CLAUDE.md`, `.claude/skills/memory-tools.md`
(which also still pointed at the gitignored `buglog.json`), and the two
`.wolf/hooks` nudges that previously told an agent to append the line
right where it was standing.

## Verification

* Every claim above read out of the code at `c1ca04d5`, not from the
documentation being replaced.
* `.wolf/*` allowlist and the `if: always()` required jobs confirmed in
`.github/workflows/ci.yml`, which is what makes the buglog-only pull
request route cheap.
* Both edited hooks pass `node --check`. No test asserts either message
string.
* `npm run lint:proof-tokens` passes. No credential appears in this
diff; variables are referenced by name only.
* No UI surface is touched, so no screenshot applies.

Note that this pull request edits two `.wolf/hooks/*.js` files, which
the CI path filter denies by executable extension, so it deliberately
runs the full required suite rather than taking the inert-path skip.

Closes #873


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified agent-engine setup, host launcher configuration,
socket-based task execution, and standalone smoke-test behavior.
* Updated environment-variable guidance, including the limited role of
the agent-engine SIF path.
* Documented bug-log workflows, merge-conflict handling, pull-request
coordination, and feature-branch requirements.
* Revised automated reminders to record feature-branch bug fixes in
pull-request details.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Aug 11, 2026
… task lifecycle

Rewrites the interaction-coverage probe so it can actually fail.

Authentication now comes from tests/e2e/support/live-auth.ts, which mints a
session through the admin one-time-token flow. It needs no password and
changes none. The previous version skipped fourteen of twenty-two controls on
the grounds that authenticating required rotating the shared demo account's
password through an admin endpoint that was returning 500s. Rotation is
forbidden, but it was never the only way in, and that reason was baked into a
committed coverage file.

Rewrites the task lifecycle assertions against what the deployment does since
PR #870. A task now progresses queued, running, done through an unprivileged
host launcher rather than resolving straight to a terminal blocked state, so
the engine-unavailable notice must be absent and the cancel button must be
offered. C17 asserts the button cancels a live task and that a second cancel
on the now-terminal task is refused with exactly 409, instead of the previous
range assertion that also passed on a 404 or a 401.

C18 no longer disables itself after its first run. It stubs an empty list
rather than requiring the live account to have no history, which earlier
controls in the same file guarantee it does.

C23 is new: it enumerates every focusable control the deployed sign-in and
tasks screens render and requires the set to equal the ledger's dom entries,
so a control cannot ship without raising the denominator. It also asserts the
absence of the transcript pane the ledger claims does not exist yet.

The coverage builder reads Playwright's own per-test verdict rather than the
first attempt of a retried test, and treats a retry-pass as unproven. It exits
non-zero when a control has no test at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Aug 13, 2026
… task lifecycle

Rewrites the interaction-coverage probe so it can actually fail.

Authentication now comes from tests/e2e/support/live-auth.ts, which mints a
session through the admin one-time-token flow. It needs no password and
changes none. The previous version skipped fourteen of twenty-two controls on
the grounds that authenticating required rotating the shared demo account's
password through an admin endpoint that was returning 500s. Rotation is
forbidden, but it was never the only way in, and that reason was baked into a
committed coverage file.

Rewrites the task lifecycle assertions against what the deployment does since
PR #870. A task now progresses queued, running, done through an unprivileged
host launcher rather than resolving straight to a terminal blocked state, so
the engine-unavailable notice must be absent and the cancel button must be
offered. C17 asserts the button cancels a live task and that a second cancel
on the now-terminal task is refused with exactly 409, instead of the previous
range assertion that also passed on a 404 or a 401.

C18 no longer disables itself after its first run. It stubs an empty list
rather than requiring the live account to have no history, which earlier
controls in the same file guarantee it does.

C23 is new: it enumerates every focusable control the deployed sign-in and
tasks screens render and requires the set to equal the ledger's dom entries,
so a control cannot ship without raising the denominator. It also asserts the
absence of the transcript pane the ledger claims does not exist yet.

The coverage builder reads Playwright's own per-test verdict rather than the
first attempt of a retried test, and treats a retry-pass as unproven. It exits
non-zero when a control has no test at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Aug 14, 2026
… task lifecycle

Rewrites the interaction-coverage probe so it can actually fail.

Authentication now comes from tests/e2e/support/live-auth.ts, which mints a
session through the admin one-time-token flow. It needs no password and
changes none. The previous version skipped fourteen of twenty-two controls on
the grounds that authenticating required rotating the shared demo account's
password through an admin endpoint that was returning 500s. Rotation is
forbidden, but it was never the only way in, and that reason was baked into a
committed coverage file.

Rewrites the task lifecycle assertions against what the deployment does since
PR #870. A task now progresses queued, running, done through an unprivileged
host launcher rather than resolving straight to a terminal blocked state, so
the engine-unavailable notice must be absent and the cancel button must be
offered. C17 asserts the button cancels a live task and that a second cancel
on the now-terminal task is refused with exactly 409, instead of the previous
range assertion that also passed on a 404 or a 401.

C18 no longer disables itself after its first run. It stubs an empty list
rather than requiring the live account to have no history, which earlier
controls in the same file guarantee it does.

C23 is new: it enumerates every focusable control the deployed sign-in and
tasks screens render and requires the set to equal the ledger's dom entries,
so a control cannot ship without raising the denominator. It also asserts the
absence of the transcript pane the ledger claims does not exist yet.

The coverage builder reads Playwright's own per-test verdict rather than the
first attempt of a retried test, and treats a retry-pass as unproven. It exits
non-zero when a control has no test at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

1 participant