Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
117 commits
Select commit Hold shift + click to select a range
64f6526
Fix export-time trust_remote_code bypass in FP8/INT8/GGUF-LoRA export…
danielhanchen Jul 5, 2026
53a071c
Harden Windows Pester install against missing PSGallery (#6892)
danielhanchen Jul 6, 2026
cb6737c
Auto Xet to HTTP download fallback in from_pretrained; share Studio's…
danielhanchen Jul 6, 2026
9407d49
GRPO: sequence packing for the no-grad old/ref logp path (default-on)…
danielhanchen Jul 6, 2026
08e133c
Add PrefixGrouper for GRPO: dedup the shared prompt across a group's …
danielhanchen Jul 6, 2026
22bd86e
Handle odd shapes and non-float scales in FP8BlockQuantLinear (#6848)
danielhanchen Jul 6, 2026
7cc1752
Scope MoE expert LoRA detection to actual MLP projection targets (#6849)
danielhanchen Jul 6, 2026
c520662
Honor an explicit sdpa or flex_attention request when flash is disabl…
danielhanchen Jul 6, 2026
c7b8666
Auto-enable grouped MoE on loaded / PEFT'd models via loader hook (#6…
danielhanchen Jul 6, 2026
95a73f0
Honor explicit load_in_16bit for local -bf16 directories (#6726)
danielhanchen Jul 6, 2026
efcaffb
Sync FORCE_FLOAT32 fallback with unsloth-zoo (gemma4, glm4_moe, qwen3…
danielhanchen Jul 6, 2026
cf4906d
Note bundled flash-linear-attention kernels for gated-deltanet models…
danielhanchen Jul 6, 2026
487b420
CI: pin lockfile-audit actions to commit SHAs (#6902)
anxkhn Jul 6, 2026
f4d1dc5
fix(fp8): use int64 offsets in weight_dequant_kernel (#6884)
anxkhn Jul 6, 2026
c44d94f
fix: map None quant method to q8_0 before lowercasing in GGUF export …
anxkhn Jul 6, 2026
cc99aab
fix: correct class name in SyntheticDataKit.chunk_data guard message …
anxkhn Jul 6, 2026
46e2cf5
studio: label RAM and VRAM readouts as GiB not GB (#6895)
danielhanchen Jul 6, 2026
2fada48
Fix llama3 RoPE scaling dropped on transformers v5 (#6907)
danielhanchen Jul 6, 2026
cb9d902
Add the second blank line before _fix_rope_inv_freq (#6910)
danielhanchen Jul 6, 2026
f0a5c52
studio: tool calling + healing parity for Llama-3, Mistral, Gemma 4 o…
danielhanchen Jul 6, 2026
f38672d
Studio: stop chat generation on the assistant-turn-end token (fixes Q…
danielhanchen Jul 6, 2026
e9f49c6
studio: deterministic backend tool-calling wiring test (#6836)
danielhanchen Jul 6, 2026
e9ea45b
Studio: coerce tool_call arguments to dict before chat templating (fi…
danielhanchen Jul 6, 2026
eb1ef44
Studio: Gemma tool-call streaming follow-ups + nested-XML escape fix …
danielhanchen Jul 6, 2026
c00c1e7
studio: tool calling for DeepSeek (R1/V3/V3.1), GLM 4.x, Kimi K2 on s…
danielhanchen Jul 6, 2026
233949c
scan_packages: baseline transitive-dep drift in the supply-chain scan…
danielhanchen Jul 7, 2026
f109e7f
Studio: parse Mistral [TOOL_CALLS] and rehearsal tool-call shapes (#5…
danielhanchen Jul 7, 2026
c2a7b78
Studio: exclude mlx-lm 0.31.3 (broke gemma4/qwen3_5 QK-norm load on A…
danielhanchen Jul 7, 2026
9dabe96
Studio chat: tool-call nudging on by default (API stays opt-in) (#6883)
danielhanchen Jul 7, 2026
8ba46b5
Studio: close switch/cancel races during model load (#6918)
danielhanchen Jul 7, 2026
46ab683
Studio: client-tool passthrough healing for safetensors and MLX (#6870)
danielhanchen Jul 7, 2026
3506371
Studio: keep the nudge wiring test collectable without the unsloth st…
danielhanchen Jul 7, 2026
9674e88
Studio: serialize the compare-mode dispatcher lifecycle to fix a star…
danielhanchen Jul 7, 2026
5608081
Studio: apply presence_penalty on the safetensors and MLX inference p…
danielhanchen Jul 7, 2026
af93868
Fix repeated base model downloads across checkpoint exports (#6896)
shimmyshimmer Jul 7, 2026
08226c2
Studio: fix torch CUDA undefined-symbol errors from a conflicting LD_…
danielhanchen Jul 7, 2026
69f8e0b
Clear stale yolo approval state on no-launch reruns (#6868)
danielhanchen Jul 7, 2026
296cacb
ROCm-on-WSL: support discrete Radeon (RDNA 3/4) in WSL, not just Stri…
LeoBorcherding Jul 7, 2026
bdb958e
Guard RoPE scaling against the transformers v5 buffer blank; honor ex…
danielhanchen Jul 7, 2026
4145037
Run the malware gate on the RAG embedding model before it loads (#6887)
danielhanchen Jul 7, 2026
d79495d
Add RDNA 2/3/4 ROCm routing tests via a CPU-only torch spoof (#6935)
danielhanchen Jul 7, 2026
59977f9
GRPO: default router_aux_loss_coef to 0 on TRL >= 1.7.0 (#6938)
danielhanchen Jul 7, 2026
411c4d1
Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6…
danielhanchen Jul 7, 2026
10d8f98
Versioning
danielhanchen Jul 7, 2026
ba450b4
Studio: add assistant response details panel (#6842)
Etherll Jul 7, 2026
8efcc17
Studio: account for DeepSeek-V4 compute buffer in context auto-fit (#…
danielhanchen Jul 7, 2026
37075c5
Bump install.sh / install.ps1 pin to unsloth>=2026.7.1 (#6943)
danielhanchen Jul 7, 2026
07ecdb3
Sort chat recents by last activity (#6844)
NilayYadav Jul 7, 2026
93c9d6d
Studio: render \[ \] and \( \) LaTeX delimiters in chat (#6914)
oobabooga Jul 7, 2026
304b8ec
fix: match qwen3-thinking double-newline in train_on_responses_only r…
InfoSage05 Jul 7, 2026
a9db53e
Studio: stream reasoning tokens in the tool-loop generator (fixes Dee…
oobabooga Jul 7, 2026
01b8085
Create ossf.yml (#6952)
danielhanchen Jul 8, 2026
49d1fb3
Speed up Studio startup path (#6899)
wasimysaid Jul 8, 2026
e7e6a0f
Polish assistant message actions menu (#6962)
shimmyshimmer Jul 8, 2026
7f9964f
Move New badge to System settings tab (#6963)
shimmyshimmer Jul 8, 2026
393d7e9
Fix opencode Unsloth provider selection (#6906)
Imagineer99 Jul 8, 2026
baacbd0
Fix Hermes install hint on Windows (#6903)
Imagineer99 Jul 8, 2026
a113f89
Studio: heal DiffusionGemma tool calls into structured tool_calls (#6…
oobabooga Jul 8, 2026
df6b5a5
Fix case-variant model matching and GGUF cache reuse in unsloth start…
Imagineer99 Jul 8, 2026
f1a2621
Studio: show Hugging Face address on hover for Hub and online model r…
danielhanchen Jul 8, 2026
de60a3a
Studio: fix currency and indentation edge cases in LaTeX rendering (#…
danielhanchen Jul 8, 2026
38dacb8
Add MLX backend support for CLI unsloth train (#6709)
Lyxot Jul 8, 2026
2a6abe2
feat(cli): support MLX distributed inference (#6845)
Lyxot Jul 8, 2026
934f879
feat(mlx): route trainer callbacks (#6929)
Lyxot Jul 8, 2026
07c8bbb
(GRPO) Fix PEFT replacement for TRL >= 1.7.0, add missing compute_aux…
marcandrelarochelle Jul 8, 2026
0e1ed88
version-compat CI: fake CPU training runs for SFT/GRPO/DPO (#6965)
danielhanchen Jul 8, 2026
6ef0936
Fix OpenClaw start default to local TUI (#6937)
Imagineer99 Jul 8, 2026
e86b787
feat: detect installed coding agent CLIs in Studio settings (#6909)
ErenAta16 Jul 8, 2026
41dd95e
Studio: don't pin transformers before the training worker activates t…
danielhanchen Jul 8, 2026
fcb1152
Studio: source CPU llama.cpp prebuilts from unslothai/llama.cpp (#6311)
oobabooga Jul 8, 2026
d0c8d55
fix(studio/hub): apply repo_id length limit per segment, not whole st…
Anai-Guo Jul 8, 2026
62a6eb2
MoE LoRA: auto-target per-expert Linear experts (gpt-oss 4bit) instea…
danielhanchen Jul 8, 2026
03cbe21
Studio: fix flash-attn and torchao install on Blackwell (sm_100+) GPU…
ThomasEricB Jul 8, 2026
38ea267
Versioning
danielhanchen Jul 8, 2026
3d41e58
Add has_blackwell_gpu to the mlx worker test's wheel_utils stub (#6980)
danielhanchen Jul 8, 2026
116ce48
Studio: allow CPU-only DiffusionGemma by granting the diffusion runne…
danielhanchen Jul 8, 2026
1a274c4
Bump install.sh / install.ps1 pins to unsloth>=2026.7.2 and unsloth-z…
danielhanchen Jul 8, 2026
5c2e536
Studio: render thinking blocks for safetensors inference with prefill…
shimmyshimmer Jul 8, 2026
7a9fb44
Remove API menu new badge (#6983)
shimmyshimmer Jul 8, 2026
92c3e48
Fix BAD_MAPPINGS not redirecting the -unsloth-bnb-4bit dynamic quants…
vineethsaivs Jul 8, 2026
dc4618c
Fix duplicate unsloth/gemma-2b-bnb-4bit mapper key routing the base 4…
vineethsaivs Jul 8, 2026
85a068c
Fix to_sharegpt optional block rendering "None" for missing extra col…
vineethsaivs Jul 8, 2026
81f789b
Guard FP8 Triton launches with tensor device context (#6888)
ramisworld Jul 8, 2026
3b73cd8
Fix per-block ID collisions and add block cleanup for unstructured up…
NilayYadav Jul 9, 2026
1b82521
Stabilize floating monitor drag (#6984)
shimmyshimmer Jul 9, 2026
8205d4c
Retry the Studio UI shutdown re-login on transient goto timeout (#7027)
danielhanchen Jul 9, 2026
5e43c62
Fix FastSentenceTransformer Qwen embedding preprocessing (#6939)
Etherll Jul 9, 2026
6d674e5
unsloth start: warn before running an agent's remote installer (#7024)
danielhanchen Jul 9, 2026
0d4bd50
Restore process-global torch.compile config on torch 2.12 so gradient…
danielhanchen Jul 9, 2026
b509d47
Silence torch._check_is_size FutureWarning and shim it if torch remov…
danielhanchen Jul 9, 2026
c1e06e9
unsloth start: add --persist to keep and reopen agent sessions (#7014)
danielhanchen Jul 9, 2026
eb775d3
Studio /v1/messages: accept thinking and unknown content blocks (#7017)
danielhanchen Jul 9, 2026
3502335
Studio: add Vulkan llama.cpp support (#5819)
oobabooga Jul 9, 2026
216a1fa
Fix Windows installer torch index override (#6972)
alkinun Jul 9, 2026
cd9d251
Fix fast inference crash on compressed-tensors FP8 models (#7025)
danielhanchen Jul 9, 2026
534c877
Keep native RoPE scaling when extending context; carry rope_theta for…
danielhanchen Jul 9, 2026
b5dca66
scripts: refresh scan_packages allowlist baseline (#7032)
danielhanchen Jul 9, 2026
fb5dc91
Studio: remove dead direct_linux_release_plan path (#7030)
danielhanchen Jul 9, 2026
d4fbc81
Restore dropped FP8 weight_scale_inv tensors on load (#6978)
danielhanchen Jul 9, 2026
b5aef63
Studio: resolve the repo-root MTP drafter after the MTP/ GGUF rename …
danielhanchen Jul 9, 2026
6a9b77e
Studio: harden OpenAI-compatible GGUF streaming (#6950)
Apoze Jul 9, 2026
86602a5
Studio: auto-load last used local model (#6966)
alkinun Jul 9, 2026
b0b8aea
Clarify in README that -H 0.0.0.0 starts a public Cloudflare tunnel (…
oobabooga Jul 10, 2026
fbcd3fa
CI: retry transient HTTP timeouts in Studio smoke probes (#7052)
danielhanchen Jul 10, 2026
33119c9
fix: guard remove_special_tokens against tokenizers without a BOS tok…
vineethsaivs Jul 10, 2026
fef37cb
Studio: queue local GGUF OpenAI-compatible requests before llama-serv…
Apoze Jul 10, 2026
7bfa209
Studio: hint at Model auto-switch in the OpenAI "No model loaded" 400…
oobabooga Jul 10, 2026
d105bd7
Studio: detect Windows Intel GPUs via the registry before WMI (#7064)
oobabooga Jul 10, 2026
29d015f
Studio: add Voice settings tab (dictation, dictionary, read aloud)
shimmyshimmer Jul 11, 2026
7c9e732
Studio: drop the single option STT engine select, rename TTS option
shimmyshimmer Jul 11, 2026
a029e3b
Studio: harden Voice settings against edge cases found in simulation
shimmyshimmer Jul 11, 2026
2ccc1ba
Studio: address Voice settings review feedback
shimmyshimmer Jul 11, 2026
830e333
Studio: use the chat mic icon in Voice settings for consistency
shimmyshimmer Jul 11, 2026
fc65256
Studio: address second round of Voice settings review feedback
shimmyshimmer Jul 12, 2026
6f838d4
Studio: drop empty and duplicate voiceURIs so the Voice tab never ren…
danielhanchen Jul 12, 2026
f33b0cf
Studio: guard dictation mic lifecycle in Voice test and Compare composer
danielhanchen Jul 12, 2026
be73912
Studio: fix dictation and read-aloud lifecycle edge cases in Voice se…
danielhanchen Jul 12, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
164 changes: 160 additions & 4 deletions .github/scripts/agent-guides-drive.sh
Original file line number Diff line number Diff line change
Expand Up @@ -154,7 +154,12 @@ raw_env() { # $1 = var name -> value (one shlex-quote layer stripped)
# writers as a side effect (it writes each agent's relocated session config).
parse_connect() {
local raw="$LOGS_DIR/connect-${AGENT}.txt"
if ! unsloth start "$AGENT" --no-launch --api-key "$UNSLOTH_API_KEY" > "$raw" 2>&1; then
# CONNECT_YOLO=1 adds --yolo. opencode/openclaw gate tool approval through their
# config (which now prompts by default), so the file-edit test opts into auto-approval
# here, the same intent as claude/codex's per-call bypass flags.
local yolo=()
[ -n "${CONNECT_YOLO:-}" ] && yolo=(--yolo)
if ! unsloth start "$AGENT" --no-launch "${yolo[@]}" --api-key "$UNSLOTH_API_KEY" > "$raw" 2>&1; then
cat_redacted "$raw"
guide_fail "'unsloth start ${AGENT} --no-launch' exited non-zero"
fi
Expand Down Expand Up @@ -371,7 +376,7 @@ case "$MODE" in
hermes) patch_hermes_tools none
invoke_via_connect "$OUT" -z "$PROMPT" ;;
openclaw) patch_openclaw_agent notools
invoke_via_connect "$OUT" agent --local --agent ci \
CONNECT_CMD_OVERRIDE=openclaw invoke_via_connect "$OUT" agent --local --agent ci \
--model "unsloth/${UNSLOTH_MODEL_ID}" --message "$PROMPT" ;;
*) invoke_via_connect "$OUT" "$PROMPT" ;;
esac
Expand All @@ -394,7 +399,10 @@ case "$MODE" in
T2='Run hello.py with python and show me the exact output.'

# The start.py recipe writers + crosscheck must see the repo; run them
# from the repo root BEFORE cd-ing into the scratch work dir.
# from the repo root BEFORE cd-ing into the scratch work dir. opencode/openclaw
# gate tool approval through their config (prompting by default), so file-edit
# opts them into auto-approval to run edits/commands headlessly.
case "$AGENT" in opencode|openclaw) CONNECT_YOLO=1 ;; esac
parse_connect
crosscheck_contract
# File-edit needs real tools, so we cannot zero them as in connection.
Expand Down Expand Up @@ -441,7 +449,7 @@ case "$MODE" in
fi ;;
opencode) invoke_via_connect "$out" run "$prompt" ;;
hermes) invoke_via_connect "$out" -z "$prompt" ;;
openclaw) invoke_via_connect "$out" agent --local --agent ci \
openclaw) CONNECT_CMD_OVERRIDE=openclaw invoke_via_connect "$out" agent --local --agent ci \
--model "unsloth/${UNSLOTH_MODEL_ID}" --message "$prompt" ;;
*) invoke_via_connect "$out" "$prompt" ;;
esac
Expand Down Expand Up @@ -519,6 +527,154 @@ case "$MODE" in
echo "[claude] attribution A/B OK (suppressed HIT, header=1 MISS)"
;;

# ── resume: does a launched agent's session survive exit and resume? ────
# Unlike the other modes, this drives the real LAUNCH path (`unsloth start
# <agent> ...`, the interactive default), not the --no-launch recipe. That
# path relocates each agent's home to a throwaway temp dir wiped on exit, so
# a session cannot be resumed -- unless --persist routes it to the stable
# Unsloth agents dir instead. We run one headless turn per pass and check
# whether the turn left a session in a persistent store (deterministic, no
# reliance on the model recalling anything), for a baseline pass and a
# --persist pass, and assert the expected split for this agent.
resume)
CODEWORD="PLATYPUS7"
T1="Remember this codeword for later: ${CODEWORD}. Reply with just the word OK."
T2="What codeword did I ask you to remember? Reply with just that word."
WORK="$WORKDIR_BASE/${AGENT}-resume"

# STABLE_HOME: the stable dir that --no-launch (and --persist) relocate to.
# Read it from a --no-launch probe (which also writes the agent's config
# there). codex/pi relocate their whole home/HOME here; opencode/claude keep
# their session data in a fixed user dir, so STABLE_HOME stays empty for them.
parse_connect
case "$AGENT" in
codex) STABLE_HOME="$(raw_env CODEX_HOME)" ;;
pi) STABLE_HOME="$(raw_env HOME)" ;;
*) STABLE_HOME="" ;;
esac

# The persistent stores a session would land in if it were NOT wiped. We
# count files here before/after each turn; a positive delta means the
# session persisted (is resumable), zero means it went to a wiped temp dir.
resume_tracked_dirs() {
case "$AGENT" in
codex) printf '%s\n' "$HOME/.codex" ;;
opencode) printf '%s\n' "$HOME/.local/share/opencode" "$HOME/.config/opencode" ;;
claude) printf '%s\n' "$HOME/.claude" ;;
pi) printf '%s\n' "$HOME/.pi" ;;
*) : ;;
esac
[ -n "$STABLE_HOME" ] && printf '%s\n' "$STABLE_HOME"
}
count_session_files() {
local total=0 d n
while IFS= read -r d; do
[ -n "$d" ] && [ -d "$d" ] || continue
n="$(find "$d" -type f 2>/dev/null | wc -l)"; total=$((total + n))
done < <(resume_tracked_dirs)
echo "$total"
}

# The headless first-turn subcommand per agent (mirrors file-edit's map),
# forwarded verbatim through the launch path as passthrough args.
set_t1_cmd() {
case "$AGENT" in
claude) T1_CMD=("${CLAUDE_CONNECT_FLAGS[@]}" -p "$T1") ;;
codex) T1_CMD=(exec "$T1") ;;
opencode) T1_CMD=(run "$T1") ;;
pi) T1_CMD=(-p "$T1") ;;
*) guide_fail "resume mode does not cover agent '$AGENT'" ;;
esac
}

# Run one headless turn through the launch path. $1=outfile, $2="" or
# "--persist", rest = the agent subcommand. --yolo auto-approves so no tool
# prompt can hang; --api-key attaches to the already-served CI model.
launch_turn() {
local out="$1" rflag="$2"; shift 2
local flag=(); [ -n "$rflag" ] && flag=("$rflag")
run_timed "$out" unsloth start "$AGENT" "${flag[@]}" --yolo \
--api-key "$UNSLOTH_API_KEY" "$@"
local rc=$?
redact "$out"
return "$rc"
}

# One pass: fresh work dir, one planting turn, set RESULT to PERSISTED/WIPED
# from the session-store delta. Runs in the main shell (not a command
# substitution) so a hang's guide_fail actually fails the job and the
# progress lines reach the CI log. $1 = "" (baseline) or "--persist".
RESULT=""
run_pass() {
local rflag="$1" label="baseline"
[ -n "$rflag" ] && label="resume"
rm -rf "$WORK"; mkdir -p "$WORK"
set_t1_cmd
local out="$LOGS_DIR/${AGENT}-resume-${label}.txt"
local before after rc
before="$(count_session_files)"
pushd "$WORK" >/dev/null || guide_fail "could not enter work dir $WORK"
launch_turn "$out" "$rflag" "${T1_CMD[@]}"; rc=$?
popd >/dev/null || true
after="$(count_session_files)"
echo "[$AGENT] ${label}: session files ${before} -> ${after} (rc=${rc})"
# The turn must succeed for the delta to mean anything: an agent that writes a
# session file then errors would otherwise be misread as PERSISTED. Mirror the
# file-edit mode and fail the pass on a non-zero launch (the flagship codex recall
# below stays WARN-only, driven by its own launch_turn calls).
[ "$rc" -eq 0 ] || { echo "[$AGENT] ${label} transcript (tail):"; tail -30 "$out" 2>/dev/null || true; \
guide_fail "resume ${label} turn for ${AGENT} exited non-zero (rc=${rc})"; }
if [ "$after" -gt "$before" ]; then RESULT="PERSISTED"; else RESULT="WIPED"; fi
}

run_pass ""; BASELINE="$RESULT"
# Only the temp-dir agents (codex/pi) need the --persist pass to prove the fix.
# opencode/claude persist either way, so the baseline already proves it and a
# second full CPU turn only risks a timeout; skip it for them.
case "$AGENT" in
codex|pi) run_pass "--persist"; RESUME="$RESULT" ;;
*) RESUME="n/a (persists either way)" ;;
esac

# Expected: codex/pi relocate their whole home to the temp dir, so a plain
# launch is WIPED and only --persist PERSISTS. opencode/claude keep their
# session data in a fixed user dir, so the baseline already PERSISTS.
case "$AGENT" in
codex|pi) EXPECT_BASELINE="WIPED" ;;
opencode|claude) EXPECT_BASELINE="PERSISTED" ;;
esac

echo "──────────────────────────────────────────────"
echo "[$AGENT] RESUME EXPERIMENT"
echo " baseline (unsloth start ${AGENT}): ${BASELINE} (expected ${EXPECT_BASELINE})"
echo " with --persist (unsloth start ${AGENT} --persist): ${RESUME}"
echo "──────────────────────────────────────────────"

[ "$BASELINE" = "$EXPECT_BASELINE" ] \
|| guide_fail "baseline resume behavior for ${AGENT} was ${BASELINE}, expected ${EXPECT_BASELINE}"
case "$AGENT" in
codex|pi)
[ "$RESUME" = "PERSISTED" ] \
|| guide_fail "--persist did not persist ${AGENT}'s session (got ${RESUME}); the session dir is still not stable" ;;
esac

# Flagship behavioral proof (codex only, WARN-only): after a --persist plant,
# resume the session and check the model actually recalls the codeword. A
# miss is not a failure (the CI model is small); the mechanism gate above is
# the real assertion.
if [ "$AGENT" = "codex" ]; then
rm -rf "$WORK"; mkdir -p "$WORK"
( cd "$WORK" && launch_turn "$LOGS_DIR/codex-resume-plant.txt" "--persist" exec "$T1" ) || true
( cd "$WORK" && launch_turn "$LOGS_DIR/codex-resume-recall.txt" "--persist" exec resume --last "$T2" ) || true
if grep -q "$CODEWORD" "$LOGS_DIR/codex-resume-recall.txt" 2>/dev/null; then
echo "[codex] behavioral recall HIT: resumed session remembered ${CODEWORD}"
else
echo "::warning::[codex] behavioral recall MISS (small CI model); mechanism gate still passed"
fi
fi
echo "[$AGENT] resume OK"
;;

*)
echo "agent-guides-drive.sh: unknown mode '$MODE'" >&2
exit 2
Expand Down
3 changes: 3 additions & 0 deletions .github/workflows/consolidated-tests-ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -364,6 +364,9 @@ jobs:
tests/utils/test_attention_masks.py \
tests/utils/test_trunc_normal_patch.py \
tests/python/test_fast_language_model_text_only.py \
tests/test_bad_mappings_redirect.py \
tests/test_prefetch_snapshot_scope.py \
tests/test_gemma_2b_mapper_key.py \
--deselect 'tests/utils/test_attention_masks.py::test_run_attention_flash_varlen_receives_window_and_softcap'
# The deselected test monkeypatches flash_attn_varlen_func, which is
# only bound on the module when `flash_attn` is importable. flash_attn
Expand Down
170 changes: 170 additions & 0 deletions .github/workflows/local-agent-guides-ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -471,6 +471,176 @@ jobs:
redacted-configs/
retention-days: 7

# ═════════════════════════════════════════════════════════════════════
# Job: resume
# Does a conversation started with `unsloth start <agent>` survive exit
# and resume? This drives the REAL launch path (not the --no-launch
# recipe the other jobs use). A plain launch relocates the agent home to
# a temp dir wiped on exit, so codex/pi cannot resume; --persist routes the
# session to the stable Unsloth agents dir so it persists. opencode/claude
# keep their session data in a fixed user dir, so they persist either way.
# Dispatch-only: it is an end-to-end experiment, not a PR gate.
# ═════════════════════════════════════════════════════════════════════
resume:
name: resume (${{ matrix.agent }})
if: github.event_name == 'workflow_dispatch'
runs-on: ubuntu-latest
timeout-minutes: 60
strategy:
fail-fast: false
matrix:
# codex/pi relocate their whole home (resume broken without --persist);
# opencode/claude keep session data in a fixed dir (resume already works).
# One agent from each class proves the split end to end; openclaw/hermes
# share codex's relocation mechanism and are covered by the unit tests.
agent: [codex, opencode, claude, pi]
env:
GGUF_REPO: unsloth/gemma-4-E4B-it-GGUF
GGUF_FILE: gemma-4-E4B-it-UD-Q4_K_XL.gguf
STUDIO_PORT: '18904'
steps:
- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
with:
persist-credentials: false

- name: Linux deps for llama.cpp prebuilt
run: |
sudo apt-get update
sudo apt-get install -y --no-install-recommends \
libcurl4-openssl-dev libssl-dev jq
- uses: actions/setup-node@48b55a011bda9f5d6aeb4c2d9c7362e8dae4041e # v6.4.0
with:
node-version: '22'

- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405 # v6.2.0
with:
python-version: '3.12'
cache: 'pip'

- name: Restore GGUF model file
id: cache-gguf
uses: actions/cache/restore@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
continue-on-error: true
with:
path: gguf-cache
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-v1

- name: Download GGUF if cache miss
id: download-gguf
if: steps.cache-gguf.outputs.cache-hit != 'true' || steps.cache-gguf.outcome != 'success'
env:
HF_TOKEN: ${{ secrets.HF_TOKEN }}
run: |
python -m pip install --upgrade huggingface_hub
mkdir -p gguf-cache
bash .github/scripts/hf-download-with-retry.sh "$GGUF_REPO" "$GGUF_FILE" gguf-cache
- name: Save GGUF model file
if: always() && steps.download-gguf.outcome == 'success'
uses: actions/cache/save@27d5ce7f107fe9357f9df03efb73ab90386fccae # v5.0.5
with:
path: gguf-cache
key: ${{ runner.os }}-gguf-${{ env.GGUF_REPO }}-${{ env.GGUF_FILE }}-v1

- name: Install Studio (--local, --no-torch)
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
HF_TOKEN: ${{ secrets.HF_TOKEN }}
run: |
mkdir -p logs
set -o pipefail
bash install.sh --local --no-torch 2>&1 | tee logs/install.log
- name: Serve unsloth run --disable-tools (gemma-4-E4B)
run: |
unsloth studio reset-password
bash .github/scripts/serve-unsloth-run.sh \
--gguf-file "$GITHUB_WORKSPACE/gguf-cache/${GGUF_FILE}" \
--port "$STUDIO_PORT" --log-dir logs \
--extra "--seed $UNSLOTH_SEED --temp 0" \
--health-timeout 900
- name: Preflight the agent's API dialect (class-a isolation)
env:
AGENT: ${{ matrix.agent }}
run: |
set -uo pipefail
B="$UNSLOTH_BASE_URL"; K="$UNSLOTH_API_KEY"
preflight_fail() {
echo "::error::[server/API regression] agent=$AGENT: $* (preflight failed BEFORE install/connect). Endpoint contract lives in studio/backend/routes/**.";
exit 1
}
code=$(curl -s -o /tmp/pf.json -w '%{http_code}' "$B/v1/models" \
-H "Authorization: Bearer $K") || true
[ "$code" = "200" ] || preflight_fail "/v1/models returned HTTP $code"
case "$AGENT" in
claude)
code=$(curl -s -o /tmp/pf.json -w '%{http_code}' "$B/v1/messages" \
-H "Authorization: Bearer $K" -H 'content-type: application/json' \
--max-time 120 \
-d "{\"model\":\"$UNSLOTH_MODEL_ID\",\"max_tokens\":16,\"messages\":[{\"role\":\"user\",\"content\":\"Hi\"}]}") || true
[ "$code" = "200" ] || preflight_fail "/v1/messages returned HTTP $code"
;;
codex)
code=$(curl -s -o /tmp/pf.json -w '%{http_code}' "$B/v1/responses" \
-H "Authorization: Bearer $K" -H 'content-type: application/json' \
--max-time 120 \
-d "{\"model\":\"$UNSLOTH_MODEL_ID\",\"input\":\"Hi\",\"max_output_tokens\":16,\"stream\":true}") || true
[ "$code" = "200" ] || preflight_fail "/v1/responses returned HTTP $code"
;;
*)
code=$(curl -s -o /tmp/pf.json -w '%{http_code}' "$B/v1/chat/completions" \
-H "Authorization: Bearer $K" -H 'content-type: application/json' \
--max-time 120 \
-d "{\"model\":\"$UNSLOTH_MODEL_ID\",\"max_tokens\":16,\"messages\":[{\"role\":\"user\",\"content\":\"Hi\"}]}") || true
[ "$code" = "200" ] || preflight_fail "/v1/chat/completions returned HTTP $code"
;;
esac
echo "preflight OK for $AGENT"
- name: Install agent CLI (class-b isolation)
env:
AGENT: ${{ matrix.agent }}
run: bash .github/scripts/agent-guides-install.sh "$AGENT"

- name: Resume experiment (launch path)
env:
AGENT: ${{ matrix.agent }}
run: bash .github/scripts/agent-guides-drive.sh resume "$AGENT"

- name: Collect server logs (debug)
if: always()
run: |
mkdir -p logs/studio-logs
cp -r "$HOME/.unsloth/studio/logs/." logs/studio-logs/ 2>/dev/null || true
if [ -n "${UNSLOTH_API_KEY:-}" ]; then
grep -rlF "$UNSLOTH_API_KEY" logs redacted-configs agent-workdir 2>/dev/null | while IFS= read -r f; do
sed -i "s#${UNSLOTH_API_KEY}#<REDACTED>#g" "$f" 2>/dev/null || true
done
fi
- name: Stop Studio
if: always()
run: |
if [ -n "${UNSLOTH_SERVER_PID:-}" ] && [ "${UNSLOTH_SERVER_PID}" != "0" ]; then
kill "${UNSLOTH_SERVER_PID}" 2>/dev/null || true
fi
sleep 2
ss -tln 2>/dev/null | grep ":${STUDIO_PORT}" || true
- name: Upload logs
if: always()
continue-on-error: true
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: resume-${{ matrix.agent }}-log
path: |
logs/
agent-workdir/
redacted-configs/
retention-days: 7

# ═════════════════════════════════════════════════════════════════════
# Job 3: prompt-cache
# (a) curl 2-turn /v1/chat/completions: assert turn-2 cached_tokens > 0
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/lockfile-audit.yml
Original file line number Diff line number Diff line change
Expand Up @@ -60,11 +60,11 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
persist-credentials: false

- uses: actions/setup-python@v5
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
with:
python-version: '3.12'

Expand Down
Loading
Loading