Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
130 commits
Select commit Hold shift + click to select a range
64f6526
Fix export-time trust_remote_code bypass in FP8/INT8/GGUF-LoRA export…
danielhanchen Jul 5, 2026
53a071c
Harden Windows Pester install against missing PSGallery (#6892)
danielhanchen Jul 6, 2026
cb6737c
Auto Xet to HTTP download fallback in from_pretrained; share Studio's…
danielhanchen Jul 6, 2026
9407d49
GRPO: sequence packing for the no-grad old/ref logp path (default-on)…
danielhanchen Jul 6, 2026
08e133c
Add PrefixGrouper for GRPO: dedup the shared prompt across a group's …
danielhanchen Jul 6, 2026
22bd86e
Handle odd shapes and non-float scales in FP8BlockQuantLinear (#6848)
danielhanchen Jul 6, 2026
7cc1752
Scope MoE expert LoRA detection to actual MLP projection targets (#6849)
danielhanchen Jul 6, 2026
c520662
Honor an explicit sdpa or flex_attention request when flash is disabl…
danielhanchen Jul 6, 2026
c7b8666
Auto-enable grouped MoE on loaded / PEFT'd models via loader hook (#6…
danielhanchen Jul 6, 2026
95a73f0
Honor explicit load_in_16bit for local -bf16 directories (#6726)
danielhanchen Jul 6, 2026
efcaffb
Sync FORCE_FLOAT32 fallback with unsloth-zoo (gemma4, glm4_moe, qwen3…
danielhanchen Jul 6, 2026
cf4906d
Note bundled flash-linear-attention kernels for gated-deltanet models…
danielhanchen Jul 6, 2026
487b420
CI: pin lockfile-audit actions to commit SHAs (#6902)
anxkhn Jul 6, 2026
f4d1dc5
fix(fp8): use int64 offsets in weight_dequant_kernel (#6884)
anxkhn Jul 6, 2026
c44d94f
fix: map None quant method to q8_0 before lowercasing in GGUF export …
anxkhn Jul 6, 2026
cc99aab
fix: correct class name in SyntheticDataKit.chunk_data guard message …
anxkhn Jul 6, 2026
46e2cf5
studio: label RAM and VRAM readouts as GiB not GB (#6895)
danielhanchen Jul 6, 2026
2fada48
Fix llama3 RoPE scaling dropped on transformers v5 (#6907)
danielhanchen Jul 6, 2026
cb9d902
Add the second blank line before _fix_rope_inv_freq (#6910)
danielhanchen Jul 6, 2026
f0a5c52
studio: tool calling + healing parity for Llama-3, Mistral, Gemma 4 o…
danielhanchen Jul 6, 2026
f38672d
Studio: stop chat generation on the assistant-turn-end token (fixes Q…
danielhanchen Jul 6, 2026
e9f49c6
studio: deterministic backend tool-calling wiring test (#6836)
danielhanchen Jul 6, 2026
e9ea45b
Studio: coerce tool_call arguments to dict before chat templating (fi…
danielhanchen Jul 6, 2026
eb1ef44
Studio: Gemma tool-call streaming follow-ups + nested-XML escape fix …
danielhanchen Jul 6, 2026
c00c1e7
studio: tool calling for DeepSeek (R1/V3/V3.1), GLM 4.x, Kimi K2 on s…
danielhanchen Jul 6, 2026
233949c
scan_packages: baseline transitive-dep drift in the supply-chain scan…
danielhanchen Jul 7, 2026
f109e7f
Studio: parse Mistral [TOOL_CALLS] and rehearsal tool-call shapes (#5…
danielhanchen Jul 7, 2026
c2a7b78
Studio: exclude mlx-lm 0.31.3 (broke gemma4/qwen3_5 QK-norm load on A…
danielhanchen Jul 7, 2026
9dabe96
Studio chat: tool-call nudging on by default (API stays opt-in) (#6883)
danielhanchen Jul 7, 2026
8ba46b5
Studio: close switch/cancel races during model load (#6918)
danielhanchen Jul 7, 2026
46ab683
Studio: client-tool passthrough healing for safetensors and MLX (#6870)
danielhanchen Jul 7, 2026
3506371
Studio: keep the nudge wiring test collectable without the unsloth st…
danielhanchen Jul 7, 2026
9674e88
Studio: serialize the compare-mode dispatcher lifecycle to fix a star…
danielhanchen Jul 7, 2026
5608081
Studio: apply presence_penalty on the safetensors and MLX inference p…
danielhanchen Jul 7, 2026
af93868
Fix repeated base model downloads across checkpoint exports (#6896)
shimmyshimmer Jul 7, 2026
08226c2
Studio: fix torch CUDA undefined-symbol errors from a conflicting LD_…
danielhanchen Jul 7, 2026
69f8e0b
Clear stale yolo approval state on no-launch reruns (#6868)
danielhanchen Jul 7, 2026
296cacb
ROCm-on-WSL: support discrete Radeon (RDNA 3/4) in WSL, not just Stri…
LeoBorcherding Jul 7, 2026
bdb958e
Guard RoPE scaling against the transformers v5 buffer blank; honor ex…
danielhanchen Jul 7, 2026
4145037
Run the malware gate on the RAG embedding model before it loads (#6887)
danielhanchen Jul 7, 2026
d79495d
Add RDNA 2/3/4 ROCm routing tests via a CPU-only torch spoof (#6935)
danielhanchen Jul 7, 2026
59977f9
GRPO: default router_aux_loss_coef to 0 on TRL >= 1.7.0 (#6938)
danielhanchen Jul 7, 2026
411c4d1
Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6…
danielhanchen Jul 7, 2026
10d8f98
Versioning
danielhanchen Jul 7, 2026
ba450b4
Studio: add assistant response details panel (#6842)
Etherll Jul 7, 2026
8efcc17
Studio: account for DeepSeek-V4 compute buffer in context auto-fit (#…
danielhanchen Jul 7, 2026
37075c5
Bump install.sh / install.ps1 pin to unsloth>=2026.7.1 (#6943)
danielhanchen Jul 7, 2026
07ecdb3
Sort chat recents by last activity (#6844)
NilayYadav Jul 7, 2026
93c9d6d
Studio: render \[ \] and \( \) LaTeX delimiters in chat (#6914)
oobabooga Jul 7, 2026
304b8ec
fix: match qwen3-thinking double-newline in train_on_responses_only r…
InfoSage05 Jul 7, 2026
a9db53e
Studio: stream reasoning tokens in the tool-loop generator (fixes Dee…
oobabooga Jul 7, 2026
01b8085
Create ossf.yml (#6952)
danielhanchen Jul 8, 2026
49d1fb3
Speed up Studio startup path (#6899)
wasimysaid Jul 8, 2026
e7e6a0f
Polish assistant message actions menu (#6962)
shimmyshimmer Jul 8, 2026
7f9964f
Move New badge to System settings tab (#6963)
shimmyshimmer Jul 8, 2026
393d7e9
Fix opencode Unsloth provider selection (#6906)
Imagineer99 Jul 8, 2026
baacbd0
Fix Hermes install hint on Windows (#6903)
Imagineer99 Jul 8, 2026
a113f89
Studio: heal DiffusionGemma tool calls into structured tool_calls (#6…
oobabooga Jul 8, 2026
df6b5a5
Fix case-variant model matching and GGUF cache reuse in unsloth start…
Imagineer99 Jul 8, 2026
f1a2621
Studio: show Hugging Face address on hover for Hub and online model r…
danielhanchen Jul 8, 2026
de60a3a
Studio: fix currency and indentation edge cases in LaTeX rendering (#…
danielhanchen Jul 8, 2026
38dacb8
Add MLX backend support for CLI unsloth train (#6709)
Lyxot Jul 8, 2026
2a6abe2
feat(cli): support MLX distributed inference (#6845)
Lyxot Jul 8, 2026
934f879
feat(mlx): route trainer callbacks (#6929)
Lyxot Jul 8, 2026
07c8bbb
(GRPO) Fix PEFT replacement for TRL >= 1.7.0, add missing compute_aux…
marcandrelarochelle Jul 8, 2026
0e1ed88
version-compat CI: fake CPU training runs for SFT/GRPO/DPO (#6965)
danielhanchen Jul 8, 2026
6ef0936
Fix OpenClaw start default to local TUI (#6937)
Imagineer99 Jul 8, 2026
e86b787
feat: detect installed coding agent CLIs in Studio settings (#6909)
ErenAta16 Jul 8, 2026
41dd95e
Studio: don't pin transformers before the training worker activates t…
danielhanchen Jul 8, 2026
fcb1152
Studio: source CPU llama.cpp prebuilts from unslothai/llama.cpp (#6311)
oobabooga Jul 8, 2026
d0c8d55
fix(studio/hub): apply repo_id length limit per segment, not whole st…
Anai-Guo Jul 8, 2026
62a6eb2
MoE LoRA: auto-target per-expert Linear experts (gpt-oss 4bit) instea…
danielhanchen Jul 8, 2026
03cbe21
Studio: fix flash-attn and torchao install on Blackwell (sm_100+) GPU…
ThomasEricB Jul 8, 2026
38ea267
Versioning
danielhanchen Jul 8, 2026
3d41e58
Add has_blackwell_gpu to the mlx worker test's wheel_utils stub (#6980)
danielhanchen Jul 8, 2026
116ce48
Studio: allow CPU-only DiffusionGemma by granting the diffusion runne…
danielhanchen Jul 8, 2026
1a274c4
Bump install.sh / install.ps1 pins to unsloth>=2026.7.2 and unsloth-z…
danielhanchen Jul 8, 2026
5c2e536
Studio: render thinking blocks for safetensors inference with prefill…
shimmyshimmer Jul 8, 2026
7a9fb44
Remove API menu new badge (#6983)
shimmyshimmer Jul 8, 2026
92c3e48
Fix BAD_MAPPINGS not redirecting the -unsloth-bnb-4bit dynamic quants…
vineethsaivs Jul 8, 2026
dc4618c
Fix duplicate unsloth/gemma-2b-bnb-4bit mapper key routing the base 4…
vineethsaivs Jul 8, 2026
85a068c
Fix to_sharegpt optional block rendering "None" for missing extra col…
vineethsaivs Jul 8, 2026
81f789b
Guard FP8 Triton launches with tensor device context (#6888)
ramisworld Jul 8, 2026
3b73cd8
Fix per-block ID collisions and add block cleanup for unstructured up…
NilayYadav Jul 9, 2026
1b82521
Stabilize floating monitor drag (#6984)
shimmyshimmer Jul 9, 2026
8205d4c
Retry the Studio UI shutdown re-login on transient goto timeout (#7027)
danielhanchen Jul 9, 2026
5e43c62
Fix FastSentenceTransformer Qwen embedding preprocessing (#6939)
Etherll Jul 9, 2026
6d674e5
unsloth start: warn before running an agent's remote installer (#7024)
danielhanchen Jul 9, 2026
0d4bd50
Restore process-global torch.compile config on torch 2.12 so gradient…
danielhanchen Jul 9, 2026
b509d47
Silence torch._check_is_size FutureWarning and shim it if torch remov…
danielhanchen Jul 9, 2026
c1e06e9
unsloth start: add --persist to keep and reopen agent sessions (#7014)
danielhanchen Jul 9, 2026
eb775d3
Studio /v1/messages: accept thinking and unknown content blocks (#7017)
danielhanchen Jul 9, 2026
3502335
Studio: add Vulkan llama.cpp support (#5819)
oobabooga Jul 9, 2026
216a1fa
Fix Windows installer torch index override (#6972)
alkinun Jul 9, 2026
cd9d251
Fix fast inference crash on compressed-tensors FP8 models (#7025)
danielhanchen Jul 9, 2026
534c877
Keep native RoPE scaling when extending context; carry rope_theta for…
danielhanchen Jul 9, 2026
b5dca66
scripts: refresh scan_packages allowlist baseline (#7032)
danielhanchen Jul 9, 2026
fb5dc91
Studio: remove dead direct_linux_release_plan path (#7030)
danielhanchen Jul 9, 2026
d4fbc81
Restore dropped FP8 weight_scale_inv tensors on load (#6978)
danielhanchen Jul 9, 2026
b5aef63
Studio: resolve the repo-root MTP drafter after the MTP/ GGUF rename …
danielhanchen Jul 9, 2026
6a9b77e
Studio: harden OpenAI-compatible GGUF streaming (#6950)
Apoze Jul 9, 2026
86602a5
Studio: auto-load last used local model (#6966)
alkinun Jul 9, 2026
b0b8aea
Clarify in README that -H 0.0.0.0 starts a public Cloudflare tunnel (…
oobabooga Jul 10, 2026
fbcd3fa
CI: retry transient HTTP timeouts in Studio smoke probes (#7052)
danielhanchen Jul 10, 2026
33119c9
fix: guard remove_special_tokens against tokenizers without a BOS tok…
vineethsaivs Jul 10, 2026
fef37cb
Studio: queue local GGUF OpenAI-compatible requests before llama-serv…
Apoze Jul 10, 2026
7bfa209
Studio: hint at Model auto-switch in the OpenAI "No model loaded" 400…
oobabooga Jul 10, 2026
d105bd7
Studio: detect Windows Intel GPUs via the registry before WMI (#7064)
oobabooga Jul 10, 2026
c3feac6
Studio: route lfm2_moe (LFM2-8B-A1B) to transformers 5.3.0 (#7040)
danielhanchen Jul 11, 2026
97161c8
Studio: route models by CONFIG_MAPPING_NAMES instead of hardcoded tab…
danielhanchen Jul 11, 2026
6412efd
Studio: auto-detect completion masking markers, stop silent full-sequ…
danielhanchen Jul 11, 2026
9fa6fd4
scripts: refresh scan_packages allowlist baseline (#7078)
danielhanchen Jul 11, 2026
f899834
DeepSeek-V4: eager attention and trainable FP8 grouped experts (#7042)
danielhanchen Jul 12, 2026
275bad1
Studio: fix the manual response-template markers that never match the…
danielhanchen Jul 12, 2026
935474c
Fix SyntheticDataKit.chunk_data emitting chunks over max_tokens (#7073)
winklemad Jul 12, 2026
2a22da9
Studio: startup loading banner and mute the benign bitsandbytes ROCm …
danielhanchen Jul 13, 2026
ca979e9
Studio: add UNSLOTH_SKIP_AUTOSTART installer flag (#7093)
danielhanchen Jul 13, 2026
9e77c1e
Studio: remove AGENTS.md and CLAUDE.md from install artifacts (#7096)
danielhanchen Jul 13, 2026
c570180
Tighten Studio instruction-file cleanup boundaries (#7097)
danielhanchen Jul 13, 2026
cc85992
Fix Studio user-message overflow for long unbroken text (#7100)
Lyxot Jul 13, 2026
2573dbd
fix(studio): use writable recipe artifact path (#7044)
Lyxot Jul 13, 2026
a337c72
Fix Studio auto-titles for reasoning models (#7098)
Lyxot Jul 13, 2026
85f5292
Studio: resync model state after a llama.cpp update unloads it (#6998)
oobabooga Jul 13, 2026
f60b982
Studio: Fix torch_dtype deprecation warning on startup and ASR load (…
oobabooga Jul 13, 2026
76d7088
Studio: Show Run button for downloaded non-GGUF models in the Model H…
oobabooga Jul 13, 2026
014d08c
Studio: install torchao Windows ROCm stub in the inference worker (#7…
oobabooga Jul 13, 2026
a5eb10a
Studio: Add rename to project chat rows (#7005)
oobabooga Jul 13, 2026
ed42702
Probe xformers support on sm_120 instead of disabling it by version (…
oobabooga Jul 14, 2026
d947638
fix(studio): prevent auth monitor reload loop
Lyxot Jul 14, 2026
c0ab90e
Add staging CI workflows for unslothai/unsloth#7118
danielhanchen Jul 14, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
164 changes: 160 additions & 4 deletions .github/scripts/agent-guides-drive.sh
Original file line number Diff line number Diff line change
Expand Up @@ -154,7 +154,12 @@ raw_env() { # $1 = var name -> value (one shlex-quote layer stripped)
# writers as a side effect (it writes each agent's relocated session config).
parse_connect() {
local raw="$LOGS_DIR/connect-${AGENT}.txt"
if ! unsloth start "$AGENT" --no-launch --api-key "$UNSLOTH_API_KEY" > "$raw" 2>&1; then
# CONNECT_YOLO=1 adds --yolo. opencode/openclaw gate tool approval through their
# config (which now prompts by default), so the file-edit test opts into auto-approval
# here, the same intent as claude/codex's per-call bypass flags.
local yolo=()
[ -n "${CONNECT_YOLO:-}" ] && yolo=(--yolo)
if ! unsloth start "$AGENT" --no-launch "${yolo[@]}" --api-key "$UNSLOTH_API_KEY" > "$raw" 2>&1; then
cat_redacted "$raw"
guide_fail "'unsloth start ${AGENT} --no-launch' exited non-zero"
fi
Expand Down Expand Up @@ -371,7 +376,7 @@ case "$MODE" in
hermes) patch_hermes_tools none
invoke_via_connect "$OUT" -z "$PROMPT" ;;
openclaw) patch_openclaw_agent notools
invoke_via_connect "$OUT" agent --local --agent ci \
CONNECT_CMD_OVERRIDE=openclaw invoke_via_connect "$OUT" agent --local --agent ci \
--model "unsloth/${UNSLOTH_MODEL_ID}" --message "$PROMPT" ;;
*) invoke_via_connect "$OUT" "$PROMPT" ;;
esac
Expand All @@ -394,7 +399,10 @@ case "$MODE" in
T2='Run hello.py with python and show me the exact output.'

# The start.py recipe writers + crosscheck must see the repo; run them
# from the repo root BEFORE cd-ing into the scratch work dir.
# from the repo root BEFORE cd-ing into the scratch work dir. opencode/openclaw
# gate tool approval through their config (prompting by default), so file-edit
# opts them into auto-approval to run edits/commands headlessly.
case "$AGENT" in opencode|openclaw) CONNECT_YOLO=1 ;; esac
parse_connect
crosscheck_contract
# File-edit needs real tools, so we cannot zero them as in connection.
Expand Down Expand Up @@ -441,7 +449,7 @@ case "$MODE" in
fi ;;
opencode) invoke_via_connect "$out" run "$prompt" ;;
hermes) invoke_via_connect "$out" -z "$prompt" ;;
openclaw) invoke_via_connect "$out" agent --local --agent ci \
openclaw) CONNECT_CMD_OVERRIDE=openclaw invoke_via_connect "$out" agent --local --agent ci \
--model "unsloth/${UNSLOTH_MODEL_ID}" --message "$prompt" ;;
*) invoke_via_connect "$out" "$prompt" ;;
esac
Expand Down Expand Up @@ -519,6 +527,154 @@ case "$MODE" in
echo "[claude] attribution A/B OK (suppressed HIT, header=1 MISS)"
;;

# ── resume: does a launched agent's session survive exit and resume? ────
# Unlike the other modes, this drives the real LAUNCH path (`unsloth start
# <agent> ...`, the interactive default), not the --no-launch recipe. That
# path relocates each agent's home to a throwaway temp dir wiped on exit, so
# a session cannot be resumed -- unless --persist routes it to the stable
# Unsloth agents dir instead. We run one headless turn per pass and check
# whether the turn left a session in a persistent store (deterministic, no
# reliance on the model recalling anything), for a baseline pass and a
# --persist pass, and assert the expected split for this agent.
resume)
CODEWORD="PLATYPUS7"
T1="Remember this codeword for later: ${CODEWORD}. Reply with just the word OK."
T2="What codeword did I ask you to remember? Reply with just that word."
WORK="$WORKDIR_BASE/${AGENT}-resume"

# STABLE_HOME: the stable dir that --no-launch (and --persist) relocate to.
# Read it from a --no-launch probe (which also writes the agent's config
# there). codex/pi relocate their whole home/HOME here; opencode/claude keep
# their session data in a fixed user dir, so STABLE_HOME stays empty for them.
parse_connect
case "$AGENT" in
codex) STABLE_HOME="$(raw_env CODEX_HOME)" ;;
pi) STABLE_HOME="$(raw_env HOME)" ;;
*) STABLE_HOME="" ;;
esac

# The persistent stores a session would land in if it were NOT wiped. We
# count files here before/after each turn; a positive delta means the
# session persisted (is resumable), zero means it went to a wiped temp dir.
resume_tracked_dirs() {
case "$AGENT" in
codex) printf '%s\n' "$HOME/.codex" ;;
opencode) printf '%s\n' "$HOME/.local/share/opencode" "$HOME/.config/opencode" ;;
claude) printf '%s\n' "$HOME/.claude" ;;
pi) printf '%s\n' "$HOME/.pi" ;;
*) : ;;
esac
[ -n "$STABLE_HOME" ] && printf '%s\n' "$STABLE_HOME"
}
count_session_files() {
local total=0 d n
while IFS= read -r d; do
[ -n "$d" ] && [ -d "$d" ] || continue
n="$(find "$d" -type f 2>/dev/null | wc -l)"; total=$((total + n))
done < <(resume_tracked_dirs)
echo "$total"
}

# The headless first-turn subcommand per agent (mirrors file-edit's map),
# forwarded verbatim through the launch path as passthrough args.
set_t1_cmd() {
case "$AGENT" in
claude) T1_CMD=("${CLAUDE_CONNECT_FLAGS[@]}" -p "$T1") ;;
codex) T1_CMD=(exec "$T1") ;;
opencode) T1_CMD=(run "$T1") ;;
pi) T1_CMD=(-p "$T1") ;;
*) guide_fail "resume mode does not cover agent '$AGENT'" ;;
esac
}

# Run one headless turn through the launch path. $1=outfile, $2="" or
# "--persist", rest = the agent subcommand. --yolo auto-approves so no tool
# prompt can hang; --api-key attaches to the already-served CI model.
launch_turn() {
local out="$1" rflag="$2"; shift 2
local flag=(); [ -n "$rflag" ] && flag=("$rflag")
run_timed "$out" unsloth start "$AGENT" "${flag[@]}" --yolo \
--api-key "$UNSLOTH_API_KEY" "$@"
local rc=$?
redact "$out"
return "$rc"
}

# One pass: fresh work dir, one planting turn, set RESULT to PERSISTED/WIPED
# from the session-store delta. Runs in the main shell (not a command
# substitution) so a hang's guide_fail actually fails the job and the
# progress lines reach the CI log. $1 = "" (baseline) or "--persist".
RESULT=""
run_pass() {
local rflag="$1" label="baseline"
[ -n "$rflag" ] && label="resume"
rm -rf "$WORK"; mkdir -p "$WORK"
set_t1_cmd
local out="$LOGS_DIR/${AGENT}-resume-${label}.txt"
local before after rc
before="$(count_session_files)"
pushd "$WORK" >/dev/null || guide_fail "could not enter work dir $WORK"
launch_turn "$out" "$rflag" "${T1_CMD[@]}"; rc=$?
popd >/dev/null || true
after="$(count_session_files)"
echo "[$AGENT] ${label}: session files ${before} -> ${after} (rc=${rc})"
# The turn must succeed for the delta to mean anything: an agent that writes a
# session file then errors would otherwise be misread as PERSISTED. Mirror the
# file-edit mode and fail the pass on a non-zero launch (the flagship codex recall
# below stays WARN-only, driven by its own launch_turn calls).
[ "$rc" -eq 0 ] || { echo "[$AGENT] ${label} transcript (tail):"; tail -30 "$out" 2>/dev/null || true; \
guide_fail "resume ${label} turn for ${AGENT} exited non-zero (rc=${rc})"; }
if [ "$after" -gt "$before" ]; then RESULT="PERSISTED"; else RESULT="WIPED"; fi
}

run_pass ""; BASELINE="$RESULT"
# Only the temp-dir agents (codex/pi) need the --persist pass to prove the fix.
# opencode/claude persist either way, so the baseline already proves it and a
# second full CPU turn only risks a timeout; skip it for them.
case "$AGENT" in
codex|pi) run_pass "--persist"; RESUME="$RESULT" ;;
*) RESUME="n/a (persists either way)" ;;
esac

# Expected: codex/pi relocate their whole home to the temp dir, so a plain
# launch is WIPED and only --persist PERSISTS. opencode/claude keep their
# session data in a fixed user dir, so the baseline already PERSISTS.
case "$AGENT" in
codex|pi) EXPECT_BASELINE="WIPED" ;;
opencode|claude) EXPECT_BASELINE="PERSISTED" ;;
esac

echo "──────────────────────────────────────────────"
echo "[$AGENT] RESUME EXPERIMENT"
echo " baseline (unsloth start ${AGENT}): ${BASELINE} (expected ${EXPECT_BASELINE})"
echo " with --persist (unsloth start ${AGENT} --persist): ${RESUME}"
echo "──────────────────────────────────────────────"

[ "$BASELINE" = "$EXPECT_BASELINE" ] \
|| guide_fail "baseline resume behavior for ${AGENT} was ${BASELINE}, expected ${EXPECT_BASELINE}"
case "$AGENT" in
codex|pi)
[ "$RESUME" = "PERSISTED" ] \
|| guide_fail "--persist did not persist ${AGENT}'s session (got ${RESUME}); the session dir is still not stable" ;;
esac

# Flagship behavioral proof (codex only, WARN-only): after a --persist plant,
# resume the session and check the model actually recalls the codeword. A
# miss is not a failure (the CI model is small); the mechanism gate above is
# the real assertion.
if [ "$AGENT" = "codex" ]; then
rm -rf "$WORK"; mkdir -p "$WORK"
( cd "$WORK" && launch_turn "$LOGS_DIR/codex-resume-plant.txt" "--persist" exec "$T1" ) || true
( cd "$WORK" && launch_turn "$LOGS_DIR/codex-resume-recall.txt" "--persist" exec resume --last "$T2" ) || true
if grep -q "$CODEWORD" "$LOGS_DIR/codex-resume-recall.txt" 2>/dev/null; then
echo "[codex] behavioral recall HIT: resumed session remembered ${CODEWORD}"
else
echo "::warning::[codex] behavioral recall MISS (small CI model); mechanism gate still passed"
fi
fi
echo "[$AGENT] resume OK"
;;

*)
echo "agent-guides-drive.sh: unknown mode '$MODE'" >&2
exit 2
Expand Down
Loading
Loading