Trajectory manager migration v4 - #2
Open
jingshenghang wants to merge 28 commits into
Open
Conversation
added 28 commits
June 8, 2026 16:13
…ectoryTree Replace slime/agent/trajectory.py (manual subagent/wipe/final segment bookkeeping) with slime/agent/trajectory_manager.py, which folds each turn into a per-session turn-node tree routed by text prefix. Sub-agent and compaction patterns now split into independent leaves automatically. Update Anthropic/OpenAI adapters and common helpers to the new record_turn / export_token_segments API, and point the coding_agent_rl example at slime.agent.trajectory_manager.
Remove vestigial bookkeeping the turn-node TrajectoryTree made redundant: * anthropic adapter: the always-empty dispatch_id plumbing in _anthropic_blocks / _build_reply (routing is now done by the tree, not by tool_use ids). * hoist the byte-identical Session dataclass and finish_session method from both adapters into common.BaseAdapter (shared session_cls + export_token_segments drain). * trajectory_manager: delete the unreferenced _starting_chains / _leaf_of_chain helpers. No behavior change; agent adapter and trajectory tests pass.
…manager-migration-v2
Bring over the four wire/manager files from trajectory-manager-migration-v2
to land the same TrajectoryManager-based anthropic adapter on this branch:
- examples/coding_agent_rl/{README,generate}.py: switch generate() to the
list[Sample] return shape from adapter.finish_session, document the env
knob SLIME_TITO_SNAPSHOT_MIN_LOSS_TOKENS.
- slime/agent/adapters/anthropic.py: absorb the wire-side scrub / mid-list
system fold / per-sid turn cap / cc title-gen skip, route through
TrajectoryManager.
- slime/agent/adapters/common.py: slim to the shared primitives still used
by the anthropic path (TurnRecord, BaseAdapter, call_sglang_generate,
shutdown_session_tasks, ok_response).
- slime/agent/trajectory_manager.py: replace the segment-based path with
the DFS routing + LCP alignment + TITO snapshot rescue implementation.
openai.py is intentionally left untouched; adapters/__init__.py drops the
OpenAIAdapter export so the package still imports under the slimmed
common.py. The OpenAI adapter and its tests do not work under this commit
and will be cleaned up in a follow-up.
Rewrite slime/agent/adapters/openai.py on top of the new
TrajectoryManager-based architecture so the Codex CLI (wire_api="chat",
v0.30.0) running inside an e2b sandbox can drive the slime SGLang
backend the same way anthropic.py drives Claude Code.
Key wire-format alignments for Codex 0.30.0 (encoded in
_build_oai_response / _stream_chat_completion):
* Emit all parallel tool_calls in a single SSE chunk -- Codex 0.30
accumulates per-index arguments fragments across chunks and would
otherwise merge them into one tool_call with concatenated args.
* wire_message.tool_calls is truncated to the first call -- Codex
silently drops the rest on echo, which would fork node_match_key.
* When tool_calls are present, wire_message.content=None and
manager_message.content="" -- Codex splits a single
assistant-with-text-and-tool_calls into two echoed messages, so we
suppress the text on the wire side to keep the echo single-shaped.
* manager_message intentionally omits reasoning_content -- Codex
strips it on echo; reasoning token ids stay in response_ids so
loss is unaffected.
Also revert Sample.rollout_id -> Sample.group_id in
trajectory_manager.py to match the upstream Sample field rename
(rollout_id is now write-only deprecated and raises on read), which is
hit at finish_session time and is a prerequisite for the openai e2e
path to run.
Verified: pytest smoke (1 SWE instance, e2b sandbox + Codex CLI ->
OpenAIAdapter -> local sglang:30000) -> rc=0, forks=0, leaves=1,
turns=39 over 5.8M tokens with 32 tokens of expected TITO drift
(reasoning text not echoed back).
…s log * TrajectoryManager owns the snapshot threshold default (1024) — drop None-passthrough from AnthropicAdapter and the hardcoded 1000 in examples/coding_agent_rl/generate.py so the single source of truth holds. * TrajectoryManager.__init__: remove dead kwargs (tokenizer, chat_template_kwargs, end_of_turn_token_id) — none were read since plan C. * FilteredAccessLogger drops HEAD heartbeats and only emits when status != 200 or elapsed > 120s — kills the web_log.py:232 spam without silencing real errors / slow handlers.
When claude-code replays a session and reformats a prior assistant message (tool_call arg ordering, whitespace), the DFS breaks at that assistant group and every reformat would spawn a new sibling subtree. Opt-in via fork_merge_max_response_tokens: if exactly one leaf assistant sibling has turn_response_ids length < threshold, collapse onto it and mark it loss_mask=0 at linearization. Sample metadata records fork_merge_masked_tokens / fork_merge_turns; a warning logs each merge. - TrajectoryManager: __init__ kwarg, Step 1.5 in append_turn, mask=0 emit in get_trajectory; revert tito_snapshot_min_loss_tokens default back to None to keep the opt-in contract. - AnthropicAdapter / OpenAIAdapter: pass-through kwarg (only forwarded when non-None); fix OpenAIAdapter erroneously passing tokenizer= to TrajectoryManager. - examples/coding_agent_rl/generate.py: parse SLIME_FORK_MERGE_MAX_RESPONSE_TOKENS env var. E2E on 20 SWE tasks with threshold=1024: 5 rewrites merged (3164 masked tokens), asst-role forks 15->6 vs no-rescue baseline.
Rescue branch was merging the rewritten turn into the sibling node's metadata but leaving sib.messages as the pre-rewrite payload. The subsequent turn replays the rewritten payload in its prompt history, DFS-fails to match the (unchanged) sibling, falls through Step 1.5 (sibling is no longer a leaf since the new turn child attached), and forks anyway — defeating the rescue. Update sib.messages to the rewritten version at rescue time. The per-turn sglang snapshot (turn_response_ids/logprobs/turn_index) stays on the original node, and get_trajectory still emits it with loss_mask=0 via the fork_merged flag. Validated end-to-end on a 20-instance SWE batch: tool→2×assistant forks dropped 6 → 0; total forks 27 → 18.
CLAUDE_CODE_ATTRIBUTION_HEADER=0 (set in examples/coding_agent_rl/sandbox.py and the e2e test runner) tells claude-code to suppress the ``x-anthropic-billing-header: cc_version=...; cch=...;`` block it otherwise prepends to the system prompt. Verified on a 56-turn e2e batch: zero requests contained the header, no scrub mutations fired. Remove _scrub_claude_code_billing_header_in_body, its regex, the call site, and the now-unused `re` import.
…nearization TrajectoryManager now uses strict exact-prefix linearization and raises on TITO drift, so the drift_fork_min_loss_tokens / fork_merge_max_response_tokens knobs are removed from both adapters. generate.py warns loudly if the corresponding env vars are still set, and stops attaching per-trajectory metadata to merged samples (revisit when dump/analysis needs it).
Add the single tolerated exception to the strict exact-prefix TrajectoryManager contract: when cc re-renders a short prior assistant message (tool_call arg order / whitespace), DFS forks at that assistant and leaves the original short turn as a standalone stub leaf -> its own Sample, diluting the trajectory's evenly-split reward. _try_merge_assistant_rewrite absorbs such a rewrite onto the existing leaf when its response is short enough (fork_merge_max_response_tokens, default 1024), demoting that node to routing-only so it contributes 0 training tokens. Wire the threshold through Anthropic/OpenAI adapters and the coding_agent_rl generate entrypoint (env SLIME_FORK_MERGE_MAX_RESPONSE_TOKENS).
…t_trajectory) 30 cases across 3 groups: routing-tree layer (message-identity forks), linearization layer (token-id drift A/B1/B2, dedup, reward split), and combined/stress (rewrite-merge, tree-fork+token-drift, deep multi-leaf, long mixed session). Semantic token vocab + reverse table for readable data; dual mode (strict assertions + human-readable tree/sample dumps).
Each case now prints [raw turns] (the source prompt_ids/response_ids decoded to names, finish_reason, logprobs presence) before [tree] and [samples], so the full data flow source->tree->samples is visible.
- 1.7 calls get_trajectory and asserts the <DRIFT> token lands in leaf 2's stripped prompt region (loss=0), proving token drift never corrupts a trained response while still being carried in the sample tokens. - get_traj wrapper snapshots the tree before get_trajectory drains the sid, so every case (incl. group 2/3) shows [tree] and [samples] together instead of <drained>.
- All group-1 cases (1.1-1.6, 1.8, 1.9, 1.10) and 3.4 now call get_trajectory and record their samples, so [samples] with token/loss alignment is shown for every case (1.10 empty-response shows 0 samples, 1.9 records both sids). - Printer always emits the [samples] header (incl. 0). - _asst_body: counter-based label->token assignment (was a hash) so distinct labels never collide and mislabel dump tokens.
The dump previously printed only the already-divided per-sample reward, so the 'reward / n_samples' averaging wasn't visible. Now the [samples] header shows the split (input / n = per-sample) and get_traj asserts the per-sample shares sum back to the input reward (the averaging invariant).
Previously cases used arbitrary input rewards (2.0/3.0/4.0) with no semantic meaning, which was confusing. Now every get_trajectory call uses reward=1.0; per-sample split varies only by sample count (1.0/N), and assertions check the even split generically instead of magic numbers.
Whitespace-only rewrite drift (e.g. cc turning 'ok' into 'ok ') was invisible in the dump, making 3.1's rewrite-merge trigger impossible to see. _vis() now shows spaces as ␣ in [raw turns] and [samples] labels.
Coverage of trajectory_manager.py rose 94%->98%. New cases: - 4.1 tools metadata attaches to first system node only - 4.2 logprobs/ids length mismatch raises - 4.3 empty prompt_messages skipped (no-op) - 4.4 default base_sample (None) - 4.5 mixed logprobs across turns (turn2 padded 0.0) - 4.6 case-B1 drift threshold boundary (d==threshold forks, d<threshold replaces)
Replace hand-derived loss_mask index arithmetic (error-prone — it was wrong twice during review) with golden string assertions. Each sample renders to a readable line where trained tokens (loss=1) are wrapped in [...] and context (loss=0 / stripped prompt) is bare, e.g. <sys> system:S </sys> <usr> user:u </usr> <gen> [r:call] [</ast>] <tul> ... Every case now pins its FULL linearized output as one human-reviewable literal, so any change to tokens, response boundary, or which tokens carry training signal is caught. Verified: stripping the [...] brackets (a loss_mask regression) fails the assertion. 36 cases pass, 98% coverage.
Rewrite TrajectoryManager.get_trajectory to tolerate TITO re-tokenization drift instead of raising. Divergence index L is classified by where it falls: prompt region -> fork; inside most-recent response span -> replace if drifted tail < threshold else fork; inside an earlier response span -> always fork. Add cross-leaf dedup so shared snapshot nodes train exactly once. Rename fork_merge_max_response_tokens -> fork_threshold_tokens across the adapters and example generate.py.
Remove from version control while keeping local copies (git rm --cached): - docs/superpowers/specs/2026-06-08-trajectory-manager-e2e-tests-design.md - tests/test_agent/test_trajectory_manager.py - tests/test_agent/test_trajectory_manager_e2e.py
Remove branch-added inline comments and docstrings in generate.py, drop the SLIME_DRIFT_FORK_MIN_LOSS_TOKENS warning block, and strip the TurnRecord docstring in adapters/common.py.
Trim the module/why docstrings to the repo's comment conventions: keep invariants, gotchas, and cross-layer contracts (cross-leaf dedup, truncated-span loss=1, sort_keys list-order, fully-masked-segment drop); drop comments that merely restate the code. Rewrite the module docstring around the append_turn / get_trajectory data flow.
Trim dead code from generate.py and the anthropic/openai/common adapters, streamline trajectory_manager linearization, and add an end-to-end trajectory manager test.
…ulting A None base_sample should never reach get_trajectory; replace the silent Sample(index=0) fallback with an assert so callers can't drop it. Update the e2e case to expect the assert and import pytest.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.