Endless Terminals Environment Integration - #24
Conversation
| terminal_backend="local", | ||
| use_dataset=True, | ||
| tasks_base_dir="", | ||
| group_size=1, |
There was a problem hiding this comment.
can you increase the group size default here
There was a problem hiding this comment.
forgot to update from testing good callout, will do
| use_dataset=True, | ||
| tasks_base_dir="", | ||
| group_size=1, | ||
| total_steps=1, |
There was a problem hiding this comment.
same for above
| task_name = item.get("task_name", "unknown") | ||
| docker_image = item.get("docker_image", self.config.default_docker_image) | ||
|
|
||
| print(f"[DEBUG] collect_trajectory START for {task_name}", flush=True) |
There was a problem hiding this comment.
can you switch print to logging?
There was a problem hiding this comment.
yep, will do this as well
|
|
||
| async def wandb_log(self, wandb_metrics: Optional[Dict] = None): | ||
| """Log Endless Terminals specific metrics to wandb.""" | ||
| if wandb_metrics is None: |
There was a problem hiding this comment.
anyway you can add some metrics?
| return 0.0 | ||
|
|
||
| async def evaluate(self): | ||
| """Periodic evaluation (optional).""" |
There was a problem hiding this comment.
can you make an eval somehow?
There was a problem hiding this comment.
i'll look into this, yeah
There was a problem hiding this comment.
Added an eval set
| "masks": node.masked_tokens, | ||
| "scores": reward, | ||
| } | ||
| if hasattr(node, "logprobs") and node.logprobs: |
There was a problem hiding this comment.
you need to include logprobs into the scored data item
There was a problem hiding this comment.
Good callout, will do this
|
|
||
| if nodes: | ||
| # Phase 2: use actual node data | ||
| node = nodes[-1] |
There was a problem hiding this comment.
I was going off this:
hermes-agent/environments/hermes_base_env.py
Line 569 in 8b54bb4
assuming that nodes[-1] is the accumulation of the full trajectory, is that not accurate here?
There was a problem hiding this comment.
No, @teknium1 needs to fix that too, here's how you use it:
There was a problem hiding this comment.
we may have multiple trajectories in the node due to how interesting agents can be, so you may need to return multiple sequences
…on export, pinned sessions, context meter Ported from ibelick/webclaw PRs NousResearch#24, #10, NousResearch#14, NousResearch#13: - Command palette (⌘K): search and switch sessions instantly - Conversation export: download as Markdown, JSON, or Plain Text - Pinned sessions: pin/unpin from context menu, shown at top of sidebar - Context meter: token usage ring in chat header with hover details - Keyboard shortcuts: ⌘K search, ⌘⇧O new session New UI primitives: autocomplete, command, input, preview-card Attachment button/preview components (composer already has built-in support)
…f light Cherry-picked from PR NousResearch#24 (clawjasper56). One-liner: respects OS dark/light preference out of the box for new users.
…agents (phase 10) The SDK landed PRs NousResearch#24/NousResearch#25/NousResearch#26 in synadia-ai/synadia-agents: - verb-first subjects (`agents.prompt.{a}.{o}.{s}`, `agents.hb.{a}.{o}.{s}`, new `agents.status.{a}.{o}.{s}`) and `metadata.protocol_version="0.3"` - pinned `_INBOX.agents` reply-inbox prefix (caller-side; no-op for us) - `name`+`session` collapsed into a single `session_name` (the 5th subject token) — `Envelope.session` and the `session=` kwarg on `AgentService` / `Agent.prompt` are gone. One service = one session_name. Package + import root rename: `natsagent` → `synadia-ai-agents`, `synadia_ai.agents`. Service-side class `Agent` → `AgentService`. Adapter changes: - Adopt single-service-per-session: rely on Hermes profile isolation for multi-session deployments instead of building an envelope.session demuxer on top of `AgentService`. The `_session_locks` dict collapses to a single `_session_lock`. - The SDK explicitly does not own NATS connections: callers build the client. Adapter calls `nats.connect(servers=...)` or `nats.connect(**sdk.load_context_options(name))` directly. - Config: `extra.name` + `extra.session_default` → required `extra.session_name`; env var `HERMES_NATS_NAME`/`HERMES_NATS_SESSION` → `HERMES_NATS_SESSION_NAME`. No migration shim — branch hadn't merged. - Lock identity rebuilt as `{agent}:{owner}:{session_name}`. Tests + docs: - conftest mock renamed `_ensure_natsagent_mock` → `_ensure_synadia_agents_mock`, installs under `sys.modules["synadia_ai.agents"]`, also stubs `nats` so the adapter's `nats.connect(...)` resolves under test. - New `mock_nats` fixture in test_nats_connect.py; concurrent-distinct- sessions test removed (v0.2-only concept); positive test added that chat_id is sourced from `settings.session_name` regardless of any stray envelope field. - design doc §1-§6/§11/§17 updated for v0.3; progress doc gains a Phase 10 decision-log entry; user-facing nats.md rewritten with verb-first subject examples, status endpoint walkthrough, and `_INBOX.agents.>` permission note. Live-verified end-to-end against `nats-server -p 4223` + `hermes-local` context + `model: anthropic/claude-haiku-4.5` over OpenRouter: real prompt streamed a real haiku reply through `agents.prompt.hermes.rene.local`, multi-turn session continuity intact, `/status` slash command dispatched through the gateway's command registry. Discovery shows `protocol_version: 0.3`. Heartbeats fire on `agents.hb.hermes.rene.local`. Status endpoint replies on `agents.status.hermes.rene.local`. NATS gateway tests: 190/190 green. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…agents (phase 10) The SDK landed PRs NousResearch#24/NousResearch#25/NousResearch#26 in synadia-ai/synadia-agents: - verb-first subjects (`agents.prompt.{a}.{o}.{s}`, `agents.hb.{a}.{o}.{s}`, new `agents.status.{a}.{o}.{s}`) and `metadata.protocol_version="0.3"` - pinned `_INBOX.agents` reply-inbox prefix (caller-side; no-op for us) - `name`+`session` collapsed into a single `session_name` (the 5th subject token) — `Envelope.session` and the `session=` kwarg on `AgentService` / `Agent.prompt` are gone. One service = one session_name. Package + import root rename: `natsagent` → `synadia-ai-agents`, `synadia_ai.agents`. Service-side class `Agent` → `AgentService`. Adapter changes: - Adopt single-service-per-session: rely on Hermes profile isolation for multi-session deployments instead of building an envelope.session demuxer on top of `AgentService`. The `_session_locks` dict collapses to a single `_session_lock`. - The SDK explicitly does not own NATS connections: callers build the client. Adapter calls `nats.connect(servers=...)` or `nats.connect(**sdk.load_context_options(name))` directly. - Config: `extra.name` + `extra.session_default` → required `extra.session_name`; env var `HERMES_NATS_NAME`/`HERMES_NATS_SESSION` → `HERMES_NATS_SESSION_NAME`. No migration shim — branch hadn't merged. - Lock identity rebuilt as `{agent}:{owner}:{session_name}`. Tests + docs: - conftest mock renamed `_ensure_natsagent_mock` → `_ensure_synadia_agents_mock`, installs under `sys.modules["synadia_ai.agents"]`, also stubs `nats` so the adapter's `nats.connect(...)` resolves under test. - New `mock_nats` fixture in test_nats_connect.py; concurrent-distinct- sessions test removed (v0.2-only concept); positive test added that chat_id is sourced from `settings.session_name` regardless of any stray envelope field. - design doc §1-§6/§11/§17 updated for v0.3; progress doc gains a Phase 10 decision-log entry; user-facing nats.md rewritten with verb-first subject examples, status endpoint walkthrough, and `_INBOX.agents.>` permission note. Live-verified end-to-end against `nats-server -p 4223` + `hermes-local` context + `model: anthropic/claude-haiku-4.5` over OpenRouter: real prompt streamed a real haiku reply through `agents.prompt.hermes.rene.local`, multi-turn session continuity intact, `/status` slash command dispatched through the gateway's command registry. Discovery shows `protocol_version: 0.3`. Heartbeats fire on `agents.hb.hermes.rene.local`. Status endpoint replies on `agents.status.hermes.rene.local`. NATS gateway tests: 190/190 green. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…agents (phase 10) The SDK landed PRs NousResearch#24/NousResearch#25/NousResearch#26 in synadia-ai/synadia-agents: - verb-first subjects (`agents.prompt.{a}.{o}.{s}`, `agents.hb.{a}.{o}.{s}`, new `agents.status.{a}.{o}.{s}`) and `metadata.protocol_version="0.3"` - pinned `_INBOX.agents` reply-inbox prefix (caller-side; no-op for us) - `name`+`session` collapsed into a single `session_name` (the 5th subject token) — `Envelope.session` and the `session=` kwarg on `AgentService` / `Agent.prompt` are gone. One service = one session_name. Package + import root rename: `natsagent` → `synadia-ai-agents`, `synadia_ai.agents`. Service-side class `Agent` → `AgentService`. Adapter changes: - Adopt single-service-per-session: rely on Hermes profile isolation for multi-session deployments instead of building an envelope.session demuxer on top of `AgentService`. The `_session_locks` dict collapses to a single `_session_lock`. - The SDK explicitly does not own NATS connections: callers build the client. Adapter calls `nats.connect(servers=...)` or `nats.connect(**sdk.load_context_options(name))` directly. - Config: `extra.name` + `extra.session_default` → required `extra.session_name`; env var `HERMES_NATS_NAME`/`HERMES_NATS_SESSION` → `HERMES_NATS_SESSION_NAME`. No migration shim — branch hadn't merged. - Lock identity rebuilt as `{agent}:{owner}:{session_name}`. Tests + docs: - conftest mock renamed `_ensure_natsagent_mock` → `_ensure_synadia_agents_mock`, installs under `sys.modules["synadia_ai.agents"]`, also stubs `nats` so the adapter's `nats.connect(...)` resolves under test. - New `mock_nats` fixture in test_nats_connect.py; concurrent-distinct- sessions test removed (v0.2-only concept); positive test added that chat_id is sourced from `settings.session_name` regardless of any stray envelope field. - design doc §1-§6/§11/§17 updated for v0.3; progress doc gains a Phase 10 decision-log entry; user-facing nats.md rewritten with verb-first subject examples, status endpoint walkthrough, and `_INBOX.agents.>` permission note. Live-verified end-to-end against `nats-server -p 4223` + `hermes-local` context + `model: anthropic/claude-haiku-4.5` over OpenRouter: real prompt streamed a real haiku reply through `agents.prompt.hermes.rene.local`, multi-turn session continuity intact, `/status` slash command dispatched through the gateway's command registry. Discovery shows `protocol_version: 0.3`. Heartbeats fire on `agents.hb.hermes.rene.local`. Status endpoint replies on `agents.status.hermes.rene.local`. NATS gateway tests: 190/190 green. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Fixes 12 remaining MEDIUM issues from the deep audit (19 total, 7 fixed in Round 12): design_agent: - NousResearch#15: add asyncio.wait_for(300s) around LLM API call to prevent infinite hangs - NousResearch#17: replace 2x hardcoded 'claude-opus-4-8' with shared DEFAULT_MODEL constant qa_agent / validate_agent: - NousResearch#20,NousResearch#22,NousResearch#23: already fixed in Round 12 (verified — dynamic timeout/threshold values used) memory.py: - NousResearch#24: frontmatter parser uses regex r'^---$' instead of str.split('---',2), preventing false splits on content containing '---' (SQL, markdown tables) - NousResearch#25: parse and preserve 'description' field from frontmatter in metadata, fixing write→load roundtrip data loss profiles.py: - NousResearch#26: ProfileConfig now frozen=True (immutable dataclass per coding standards) deploy_agent: - NousResearch#31: replace 2x sync subprocess.run with asyncio.create_subprocess_exec - fix 5x .decode() → .decode('utf-8', errors='replace') for Windows CJK safety - remove unused import subprocess db.py: - NousResearch#27: add class docstring explaining RLock + _unlocked pattern - NousResearch#28: FK constraints already in DDL (verified PRAGMA foreign_keys=ON active) - NousResearch#29: add _ensure_connection() with PRAGMA integrity_check(1) + auto-reconnect on 4 critical methods (create_task, get_task, claim_task, submit_result) - extract _create_connection() static method for reuse by reconnect Tests: 79 passed, 0 failed
… A/B gate; scored value-saturation
Targets the one measured weakness — within-task ranking (per-prompt ρ≈0.34) — by adding
comparative (pairwise) Δ/stakes elicitation as an OFF-BY-DEFAULT, A/B-gated experiment that
cannot regress the live skill.
Phase B (comparative elicitation):
- scripts/pairwise.py (new, pure): Bradley-Terry MLE (phantom-regularized) + win-count fallback
+ anchored [0,1] mapping. Two virtual anchors (FLOOR=no-change->0, CEILING=completely-
different->1) sit in every question's comparison set so BETWEEN-task scale is preserved while
within-question ordering is fixed.
- pipeline.judge_plan_change_pairwise[_batch]: same per-answer delta_plan/stakes contract as the
absolute judge (drop-in for voi.evsi/score_record); 2 calls/question; safe-zeroes on failure.
- infogain: value_judge_mode selector ("absolute"|"pairwise", default absolute). Absent key ->
absolute, so every cfg from DEFAULTS is byte-identical; one call site branches.
- validate_evsi --ab / --elicit-model: scores BOTH methods on one shared question/answer set
(realized measured once); analyze_evsi prints per-method within-task ρ + adopt/keep verdict.
Phase A (confirmation, eval-only):
- saturation_scan.py --scored: full-pipeline value saturation (max_value + #>=floor per breadth).
Results (host, local judge):
- Saturation: max(value) plateaus at breadth ~2 while distinct-target coverage keeps climbing ->
confirms breadth is bounded by value not coverage (modest breadth + families is right).
- A/B gate (6 prompts): pairwise non-inferior (within-task ρ ties on realized_change, +0.040 on
realized_regret) but NOT a clear win (n=6; local realized judge saturates 39% at 1.0).
Verdict: KEEP absolute as default; pairwise stays built+off for a stronger-judge re-test.
Safety invariant: default unchanged; 64 ranker + 16 investigator tests green (+11 new:
aggregator, pairwise judge, selector default-routing).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…rect the gate metric The NousResearch#24 comparative-elicitation A/B is settled with power. The first (n=6) read was corrected twice by evidence: 1. Adversarial verification refuted the "realized judge saturates" story (saturation runs backwards to signal), and showed the gate was ranking within-task against a noisy target. 2. The powered 12-prompt A/B then showed the n=6 sub-narratives were themselves small-sample noise: realized_change is NOT within-task-dead (ρ +0.30 at n=12, was +0.04 at n=6), and pairwise does not edge ahead (−0.02, was +0.07). Binding limit was power, as predicted. Gate metric (analyze_evsi.py): - by_question now aggregates realized_stakes + mean_stakes. - ab_within_task ranks WITHIN-TASK on realized_regret (realized EVSI = the thing q_value predicts), reports realized_stakes/realized_change alongside, and shows the per-prompt paired Δρ with a conservative broad-win guard (a 1-2-outlier mean can't pass). - p1c ablation gains a stakes-only formula. Powered verdict (12 prompts / 72 questions / 216 pairs per arm, local judge fixed both arms): - KEEP absolute — pairwise is slightly WORSE on every realized target (regret abs +0.360 vs pw +0.204, loses 9/12). NousResearch#24 closed as a documented negative result; pairwise stays built+off. - Do NOT build the comparative realized judge (pointless — pairwise doesn't help on projected). - Strong positive: p1c vs realized_regret ranks √(U·EVSI) BEST (+0.360) above every component (U-only +0.264, EVSI-only +0.202, stakes-only +0.157, max-Δ +0.075) → within-task ranking is modest-but-real (ρ≈0.36 ≈ original 0.34); the FROZEN formula is validated within-task too. Record corrected across findings/design-decisions/README. 78 pure/mocked + 16 investigator tests green (+ new analyze_evsi gate/aggregate tests). Live default unchanged (absolute). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…esearch#23 stays off, graded judge rejected) + eval plumbing + grouped test suites - Eval-harness families plumbing: infogain.families_cfg(), --families/--premortem arms on score_scan/validate_evsi/run_evals, per_lens() + selection_policies() in analyze_evsi, lens-tagged rows; score_scan --include-life pool fix (default now BANK-only). - NousResearch#25 pre-mortem lens validated at BOTH ladder tiers: tier-1 projected two-arm (14 cells) and tier-2 realized two-arm (6 prompts x off/on, 336 rows) — top lens by realized_regret (0.416; failure-surface 0.602 vs 0.386 others; forced-on read-only self-prunes). Auto-on confirmed, rollback untripped. Gate false positive fixed: artifact nouns removed, word-boundary hint matching for both vantage+premortem gates ("repo" != "report", "prod" != "product", old '"db "' end-of-text miss fixed). - NousResearch#23 selection policies: analyze_evsi.selection_policies verdict — every q_value policy within ~0.03 of size-matched random within-task; rel_keep_frac stays off. - Graded change judge (opt-in --graded-change-judge) + --keep-responses + evals/rejudge.py offline instrument A/B: REJECTED (anchor-clustering, q_value link 0.60->0.38); original judge stays; harness remains for future instrument tests. - --families-model / INFOGAIN_FAMILIES_MODEL override (families layer no longer hard-pinned to glm for evals); SKILL.md drift fixes (stage 1 = plan_model; --value-judge-mode documented). - Grouped test suites: tests/run.py (basic DEFAULT = mocked/offline ~1s; live opt-in via INFOGAIN_TEST_LIVE; all = 107 tests). Docs in evals/README.md. - roadmap.md reconciled (NousResearch#23 status, NousResearch#24 CLOSED outcome, wrapper DONE); findings + design docs updated with all three verdicts. Version 1.0.0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
…UB不加载SOUL)
- P045 部分回退: 删除 load_soul_identity=True,SOUL.md 不再注入子代理
- P064 部分回退: 删除 3 处 needle replace 矫正逻辑(死代码)
- conversation_loop.py 两处(~900/~1620)
- delegate_tool.py 一处(~1634)
- 保留 P064 正向价值: agent_name identity_line (You are {agent_name})
- SUBAGENT_PROTOCOL.md 注入保留(唯一治理文件注入)
- SUB 已重构为自包含(价值观+交互护栏内嵌,无 SOUL/HERMES 悬空引用)
- verify-patches.sh NousResearch#24/NousResearch#46 已同步更新
This PR covers the first implementation of endless-terminals dataset repo paper
To run this, simply point your atropos environment to the
./environments/endless_terminals/endless_terminals_env.pyenvironment.Download Tasks
Testing the environment