feat(v1): support Harbor's separate verifier environments - #2067
Conversation
ApprovabilityVerdict: Needs human review This PR introduces a significant new feature (separate verifier environments) with complex artifact collection, runtime lifecycle management, and scoring flow changes. An unresolved P1 comment identifies a bug where the tar archive initialization uses invalid bytes that would cause separate-mode scoring to fail. You can customize Macroscope's approvability policy. Learn more. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 662b48db5b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 61beb0b5b6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e8c5eeb62b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 789a40d02e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b7caabea43
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 48ed242ed9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e83057ab44
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0d20fa0fc8
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 845cd5033b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: adff2dcc3b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
This comment was marked as resolved.
This comment was marked as resolved.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 75e5f1ed62
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cccd191309
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
cccd191 to
da81097
Compare
da81097 to
e5fb459
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e5fb45997e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
e5fb459 to
a0f7c04
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a0f7c04b1c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
a0f7c04 to
fb5b386
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fb5b386b7b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
fb5b386 to
a1e275d
Compare
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit a1e275d. Configure here.
a1e275d to
b21138c
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a1e275d643
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
superseded by #2144 |
Harbor tasks can declare that their verifier runs in its own container (`[verifier].environment_mode`, `[verifier.environment]`). We ignored the first and rejected the second, so a task that asked for isolation was graded in the box the agent had just had root in. Harbor resolves the verifier's environment as `[verifier.environment]` if declared, else a deep copy of `[environment]`. That gives three cases, and only the third is a problem: - mode-only separate → the task's own image, so a second box is all it takes; - a declared environment with a `docker_image` → pull that instead; - a declared environment *without* one → Harbor builds the verifier image from `tests/Dockerfile`. Verifiers pulls and never builds, so this is rejected with a message saying to build and push the image and name the ref. `ignore_dockerfile` remains the one escape hatch, and now warns, because falling back to the agent's image runs the verifier somewhere the task never declared. Only declared artifacts cross over, which is Harbor's own model (`Trial._run_separate_verifier` uploads artifacts and nothing else) — so #2144's `collect`/`restore` and its 32 MB budget carry this unchanged. Artifacts are restored before `tests/` is staged, and `/tests` is wiped first: an artifact entry pointing into `/tests` would otherwise hand the agent the grader's own scripts, and a fresh container of the task image can ship a stale `/tests` of its own. `[[verifier.collect]]` keeps working here. The hooks run in `finalize`, in the agent's box, producing exactly the files that then travel — the two compose, so unlike #2067 there is no need to reject them. Two smaller changes: - The verifier's declared network mode is enforced rather than warned about, via the `allow`/`block` policy the runtimes already have. `allow_internet = false` means the grader really has no network, which is the point of declaring it. - A failed `tests/` staging now raises instead of leaving `test.sh` missing and scoring 0. A false zero reads as a failed attempt. `--taskset.ignore-separate-verifier` forces shared grading when a sandbox per task is too expensive. Design and prior art: xeophon in #2067, rewritten against #2144. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Harbor tasks can declare that their verifier runs in its own container (`[verifier].environment_mode`, `[verifier.environment]`). We ignored the first and rejected the second, so a task that asked for isolation was graded in the box the agent had just had root in. Harbor resolves the verifier's environment as `[verifier.environment]` if declared, else a deep copy of `[environment]`. That gives three cases, and only the third is a problem: - mode-only separate → the task's own image, so a second box is all it takes; - a declared environment with a `docker_image` → pull that instead; - a declared environment *without* one → Harbor builds the verifier image from `tests/Dockerfile`. Verifiers pulls and never builds, so this is rejected with a message saying to build and push the image and name the ref. `ignore_dockerfile` remains the one escape hatch, and now warns, because falling back to the agent's image runs the verifier somewhere the task never declared. Only declared artifacts cross over, which is Harbor's own model (`Trial._run_separate_verifier` uploads artifacts and nothing else) — so #2144's `collect`/`restore` and its 32 MB budget carry this unchanged. Artifacts are restored before `tests/` is staged, and `/tests` is wiped first: an artifact entry pointing into `/tests` would otherwise hand the agent the grader's own scripts, and a fresh container of the task image can ship a stale `/tests` of its own. `[[verifier.collect]]` keeps working here. The hooks run in `finalize`, in the agent's box, producing exactly the files that then travel — the two compose, so unlike #2067 there is no need to reject them. Two smaller changes: - The verifier's declared network mode is enforced rather than warned about, via the `allow`/`block` policy the runtimes already have. `allow_internet = false` means the grader really has no network, which is the point of declaring it. - A failed `tests/` staging now raises instead of leaving `test.sh` missing and scoring 0. A false zero reads as a failed attempt. `--taskset.ignore-separate-verifier` forces shared grading when a sandbox per task is too expensive. Design and prior art: xeophon in #2067, rewritten against #2144. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Overview
Adds Harbor v1 support for grading in a fresh verifier runtime while keeping shared-mode behavior unchanged. Separate grading uses the current rollout lifecycle: rollout-scoped servers close first, harness metrics finish on the agent runtime, and task scoring moves to an isolated runtime only after confirmed teardown.
Details
public/no-networkpolicy.pre_artifacts.sh, then transfers only declared main-service filesystem artifacts at their original paths with relative-path and exclude-pattern support./tests, and restores artifacts into the fresh runtime.reward.jsonscalars or maps withreward.txtfallback; shared-mode reward behavior remains unchanged.This keeps artifact handoff and runtime ownership explicit while using Verifiers' existing task, rollout, and server lifecycle interfaces.
Fixes RES-1090
Note
Add separate verifier runtime support for Harbor tasks
task.toml._collect_artifacts(), the agent sandbox is confirmed stopped, then a fresh verifier runtime is provisioned with artifacts and tests restored before scoring runs.Task.scoring_runtime()as an extension point,provision_runtime()context manager in runtimes/init.py, andRuntime.stop_confirmed()/teardown_confirmed()for explicit deletion confirmation across Docker and Prime runtimes.Runtime.read()now supports bounded reads with amax_byteslimit; oversized files raiseSandboxError. Reward files (reward.json) are read with this bound.Macroscope summarized b21138c.
Note
High Risk
Changes rollout teardown/scoring order, confirmed sandbox deletion, and Harbor artifact transfer with size/path limits—mistakes could leak agent state into grading or fail scoring on valid tasks.
Overview
Harbor separate verifiers now grade in a fresh runtime after the agent phase: declared artifacts are collected and validated, the agent box is torn down with confirmed deletion (
stop_confirmedon Docker/Prime), then tests and restored files run in a new verifier runtime derived from the task’s Docker/Prime policy (image, resources, workdir, public/no-network).Rollouts gain
Task.scoring_runtime(): harness scoring stays on the agent runtime; task scoring can enter a separate async context. Separate scoring is rejected for borrowed runtimes or shared tool servers.Runtime plumbing:
provision_runtime()replaces manual start/stop in agent provisioning and MCP launch;Runtime.read()accepts optionalmax_bytes(Docker, Prime streaming, default shell path); harness cleanup can run whenstop_failedis set.Harbor parsing uses Harbor’s validated config for shared vs separate mode, artifacts, verifier env, and explicit errors for allowlists, sidecars, and multi-step tasks; rewards read bounded
reward.json(scalar or map) withreward.txtfallback. Docs note separate verifier mode and updated parity gaps.Reviewed by Cursor Bugbot for commit b21138c. Bugbot is set up for automated code reviews on this repo. Configure here.
Fixes RES-1093