Skip to content

Outputs builder - #792

Merged
mikasenghaas merged 6 commits into
overhaul-results-savingfrom
outputs-builder
Jan 28, 2026
Merged

Outputs builder#792
mikasenghaas merged 6 commits into
overhaul-results-savingfrom
outputs-builder

Conversation

@willccbb

@willccbb willccbb commented Jan 28, 2026

Copy link
Copy Markdown
Member

Description

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Documentation update
  • Test improvement

Testing

  • All existing tests pass when running uv run pytest locally.
  • New tests have been added to cover the changes

Checklist

  • My code follows the style guidelines of this project as outlined in AGENTS.md
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

Additional Notes


Note

Shifts generation outputs from raw State objects to serialized RolloutOutput dicts and introduces an incremental results builder for efficient saving and display.

  • Add RolloutOutput type and update GenerateOutputs to use outputs (replacing states); update docs and all call sites/tests accordingly
  • Introduce GenerateOutputsBuilder to serialize once, stream progress, compute metadata, and support intermediate saves; final results sorted by example_id
  • Add state_to_output/states_to_outputs, remove sanitize_states; validate state_columns are JSON-serializable; errors serialized as strings
  • Update Environment.generate() to build outputs incrementally and adjust progress callbacks to use RolloutOutput
  • Modify utils (eval_utils, eval_display, logging_utils, save_utils) to consume serialized outputs; make_dataset now builds directly from outputs
  • Adapt integrations (gepa.adapter, RL trainer orchestrator) to the new outputs schema
  • Update Gym/eval/CLI tests and fixtures to the new API

Written by Cursor Bugbot for commit a17a1d1. This will update automatically on new commits. Configure here.

Comment thread verifiers/gepa/adapter.py Outdated
Comment thread verifiers/utils/save_utils.py
Comment thread verifiers/utils/eval_display.py
Comment thread verifiers/gepa/adapter.py
Comment thread verifiers/utils/eval_utils.py Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Bugbot Autofix is OFF. To automatically fix reported issues with Cloud Agents, enable Autofix in the Cursor dashboard.

Comment thread docs/reference.md
@willccbb
willccbb requested a review from mikasenghaas January 28, 2026 03:01

@mikasenghaas mikasenghaas left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lfgtm!

Comment thread docs/reference.md
### RolloutOutput

```python
class RolloutOutput(dict):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah yea funnily enough i had a type very similar to this one initially but then scratched it bc it felt too redundant. basiclly RolloutOutput = State - untyped columns + state columns and so I wasn't sure this justified an entirely new type (that now also has to be synced)

@mikasenghaas
mikasenghaas merged commit a822276 into overhaul-results-saving Jan 28, 2026
3 checks passed
hallerite added a commit that referenced this pull request Sep 2, 2026
prime-envs #792 snapshots image-supplied untracked files in four SWE
tasksets and passes them to capture_patch(ignore=...).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
xeophon pushed a commit that referenced this pull request Sep 2, 2026
Pure deletions. Every item has zero references in verifiers, prime-rl,
prime-envs and community-environments (grep + GitHub code search);
nothing here changes behaviour.

**v0 config trees (−820 lines).** `configs/eval/`, `configs/gepa/`,
`configs/rl/`, `configs/zero3.yaml`, `configs/endpoints.toml` are the
`vf-eval` / DeepSpeed formats of the removed v0 stack.
`tests/v1/test_configs.py` already excluded them; the filter goes with
them.

**pyproject.** The `[[tool.uv.dependency-metadata]] nemo-gym` block is
inert (nemo-gym stopped being a dependency in #2476; the PEP 723 header
in `tasksets/nemo_gym/server.py` is the one that is used), so `uv.lock`
loses its manifest entry. Markers `integration`, `slow`, `unit`,
`parsers`, `rubrics`, `environments` are applied by no test; the
`cellpylib` warning filter targets a package that is not in the lock.

**Dead code.**
- `Harness.run()`: the pre-session launch/resume dispatch, superseded by
`HarnessSession.turn`.
- `Agents.__iter__`/`__len__`: no reader anywhere; every env addresses
agents by attribute.
- `saved_config_path`: the fallback to the pre-#2429
`configs/<cli>.json` location. Only `configs/resolved/` is written.
- `eval`/`gepa` usage gate: the `--taskset.`/`--harness.` prefixes kept
alive for flags removed in #2157/#2237.
- `ServerBase.EXTRAS` and the `RUNTIME_PYTHON` hook in `mcp/launch.py`:
declared by no server anywhere.
- `PrimeAgentHarness.SUPPORTS_RESUME`: `ACPHarness` overrides
`session()`, so both sites that read the flag short-circuit before
reaching it.

`snapshot_untracked` was removed in the first revision and restored in
the second: prime-envs #792 calls it from four SWE tasksets
(`capture_patch(ignore=...)`). Open PRs in the env repos are now part of
the check.

Deliberately not in this PR: anything on the `Trace`/`Episode` record
model (`Episode` aggregates, `PolicySpan.drift`, `Trace.last_message`) —
those have downstream consumers outside this repo; `agent_config_fields`
(its `_declared_agent_configs` twin checks field defaults, not values,
so swapping is not a pure deletion), `WireTrace`, `InterceptionError`,
the `Branch` trainer-side properties, and every docs/docstring wording
fix. Those are separate PRs.

Checks run on the branch: `ruff check`, `ruff format --check`, `ty check
verifiers`, and `pytest tests/v1 -m "not e2e"` (82 passed).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- CURSOR_SUMMARY -->
---

> [!NOTE]
> **Low Risk**
> Mostly deletions of unreferenced configs and code; the only behavioral
tweaks are stricter CLI axis detection, resolved-config-only replay
paths, and always running full MCP sandbox installs.
> 
> **Overview**
> Removes **v0-era configuration** that is no longer part of the v1
stack: the entire `configs/eval/`, `configs/gepa/`, `configs/rl/` trees,
`configs/endpoints.toml`, and `configs/zero3.yaml`. Config validation in
`tests/v1/test_configs.py` no longer special-cases `endpoints.toml`.
> 
> **Tooling cleanup** drops the unused `nemo-gym`
`[[tool.uv.dependency-metadata]]` block from `pyproject.toml` (and the
matching `uv.lock` manifest), unused pytest markers, and a `cellpylib`
warning filter for a package not in the lock.
> 
> **Dead v1 code paths** are deleted without replacing callers:
`Harness.run()` (segment dispatch now lives on `HarnessSession.turn`),
`Agents.__iter__`/`__len__`, legacy resolved-config lookup in
`saved_config_path`, narrowed `eval`/`gepa` CLI usage gates (no
`--taskset.` / `--harness.` shortcuts), MCP sandbox install no longer
honors `ServerBase.EXTRAS` or `RUNTIME_PYTHON`, and
`PrimeAgentHarness.SUPPORTS_RESUME` (unused because ACP harnesses use
`session()`).
> 
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit
8389294. Bugbot is set up for automated
code reviews on this repo. Configure
[here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->

<!-- Macroscope's pull request summary starts here -->
<!-- Macroscope will only edit the content between these invisible
markers, and the markers themselves will not be visible in the GitHub
rendered markdown. -->
<!-- If you delete either of the start / end markers from your PR's
description, Macroscope will append its summary at the bottom of the
description. -->
> [!NOTE]
> ### Remove v0 config files, dead code, and fix CLI usage-gate bypass
in v1
> - Deletes leftover v0 configs under `configs/` (endpoint registry,
eval/gepa/rl TOMLs, `zero3.yaml`) and removes unused pytest markers,
dependency metadata, and dead class members (`Agents.__iter__`,
`Agents.__len__`, `ServerBase.EXTRAS`,
`PrimeAgentHarness.SUPPORTS_RESUME`).
> - Restricts the typed-axis usage-gate bypass in
[main.py](https://github.com/PrimeIntellect-ai/verifiers/pull/2496/files#diff-1335625d1379d7f1a4bf6a2f961d4b5fde5f3a9bfb8eb4cc736859f08698a8c8)
and
[gepa.py](https://github.com/PrimeIntellect-ai/verifiers/pull/2496/files#diff-dc87040f272d5f7871025c1cd9da91e56670d1cf1deff808037cb5bc356376bc)
to `--env.`/`--serve.` prefixes only; `--taskset.` and `--harness.`
arguments no longer suppress the usage message.
> - Removes the `configs/` fallback in
[output.py](https://github.com/PrimeIntellect-ai/verifiers/pull/2496/files#diff-07fce45dd0fa633f34d25e19962ab801b705a509774833de0d81bfe872787e8b):
`saved_config_path` now searches only `configs/resolved/`.
> - Changes non-subprocess MCP server launch in
[launch.py](https://github.com/PrimeIntellect-ai/verifiers/pull/2496/files#diff-989d87123d40d48511f9e1df74594d2e30776e9a99c11eca49d9cc20bf9e2230)
to always install a sandbox venv and use its Python, dropping the
`RUNTIME_PYTHON` override and server-class extras suffix.
> - Risk: non-subprocess MCP servers no longer honor a prebuilt
`RUNTIME_PYTHON` executable; `saved_config_path` returns `None` for runs
with JSON only under `configs/` rather than `configs/resolved/`.
>
> <!-- Macroscope's review summary starts here -->
>
> <sup><a href="https://app.macroscope.com">Macroscope</a> summarized
8389294.</sup>
> <!-- Macroscope's review summary ends here -->
>
<!-- Macroscope's pull request summary ends here -->

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants