Skip to content

PR #4946 staging CI - #2

Closed
Etherll wants to merge 22 commits into
mainfrom
pr-4946-ci
Closed

Etherll wants to merge 22 commits into
mainfrom
pr-4946-ci

Conversation

@Etherll

@Etherll Etherll commented Jun 22, 2026

Copy link
Copy Markdown
Owner

Staging-only CI validation for unslothai#4946.

This PR exists only to run fork GitHub Actions checks and should be closed after validation.

sanatb187 and others added 22 commits April 10, 2026 02:16
Reject dataset_streaming at the API boundary when hf_dataset is empty,
the dataset is vision/audio, or max_steps is not set. Probe eval split
with get_dataset_split_names before the streaming load so typos fail
immediately instead of mid-training. Guard column_names=None after map
on iterables. Hide the UI toggle for non-text configurations and clear
the stale flag when config becomes incompatible.
…emplate/format support (WIP)

Work-in-progress on top of feat/studio-dataset-streaming-mode (PR unslothai#4946):
- new test_training_streaming.py and iterable.py dataset helper
- streaming support in chat_templates.py and format_conversion.py
- additional streaming guards in trainer.py / models / routes
- frontend streaming wiring in params-section and training-config-store

Committed to preserve uncommitted work before merging latest main.
Resolved 4 conflicts from main's CPT/raw-text feature meeting the streaming feature:
- trainer.py: keep both imports (detect_streaming_dataset + prepare_raw_text_dataset);
  keep Optional[int] slice typing and add main's is_cpt param
- models/training.py: keep both _validate_dataset_slice and main's field_validators
- dataset-section.tsx: drop duplicate barrel imports; keep deep hasSeparateStreamingEvalSplit import
- training-config-store.ts: un-fuse two functions' shared tails; merge two version<10 migrations
BLOCKER: streaming + raw-text/CPT crashed on len(IterableDataset). Guard it in the
start route (reject format_type=="raw" or training_type=="Continued Pretraining")
and in isStreamingSupported (datasetFormat !== "raw").

Also:
- models/training.py: validate hf_dataset/subset/split (charset+length, block ..//);
  cap dataset slice indices (le=1e9); note validator ordering
- chat_templates.py: guard _apply_custom_mapping .map() for streaming
- trainer.py: warn when packing+streaming
- training-config-store.ts: persist-migration bump to v11 (standalone datasetStreaming
  backfill); add isVisionModel to NON_PERSISTED; toast on silent streamingCompatiblePatch
  mutations in the 4 indirect setters
- tests: route rejections (max_steps, raw/cpt), slice cap, unsafe hf_dataset
Re-merged after main fast-forwarded. Resolved 18 conflicts across 6 files,
preserving the streaming feature + review fixes while folding in main's evolution
(S3 dataset support, cache-safe load_dataset, base-VLM preflight, manual-slice
optimization, improved train-on-responses safety net):

- trainer.py (10): combine streaming load path with main's manual-slice fetch;
  keep IterableDataset import + main's load_dataset_cache_safe; keep both main's
  _chat_template_renders_empty/_preflight_first_batch and the streaming-aware
  _train_worker; streaming-skip the post-filter length check; streaming-aware
  total_steps that also prefers the trainer's processed dataset length
- models/training.py: keep dataset_streaming field alongside main's reformat
- routes/training.py: keep streaming validation (incl. raw/CPT BLOCKER guard) +
  main's deliberate-rejection passthrough comment
- chat_templates.py / format_conversion.py: keep centralized is_streaming_dataset
  helper (covers HF + torch) over main's torch-only inline detection
- training-config-store.ts: NON_PERSISTED keeps both isVisionModel + s3Config

Also fixed a non-conflict auto-merge artifact: duplicate isVisionModel/isAudioModel
bindings in dataset-section.tsx. Backend tests: 21/21 pass.
- raw_text: keep the lazy filter but skip len()-based row counting for
  IterableDatasets so raw-text / CPT can stream; guard the eval-size log
- routes/trainer: drop the raw/CPT streaming block; add a defensive
  not-streaming guard on the eval auto-split (train_test_split)
- dataset-section: streaming toggle is visible-but-disabled and lists the
  exact unmet requirement(s) in its tooltip; block embedding models
- training-start-overlay: show "streaming (no full download)" instead of a
  stuck download bar for streaming runs
- trim the streaming test suite to the high-value cases
…plit, rehydrate timing)

- routes: reject dataset_streaming for embedding training and on Apple Silicon
  (MLX); both loaders materialize the full dataset instead of streaming
- trainer: validate the base eval split name so streaming eval accepts HF slice
  syntax such as "validation[:1000]"
- training-config-store: defer the onRehydrateStorage setState to a microtask so
  it doesn't hit the store's TDZ during synchronous hydration
- test: streaming start rejects embedding models
…ty/eval bounds, gating)

Address a deeper streaming review:
- raw_text: resolve_column_names() guards IterableDataset.column_names=None
  (from_generator / unresolved features) so raw-text and CPT streaming no longer
  raise TypeError before training
- models/routes: reject HF slice syntax in train_split/eval_split when streaming
  (load_dataset(streaming=True) raises "Bad split"); reject mixed sources
  (local/S3) and embedding/MLX streaming at the API, not just in the UI
- trainer: an empty post-slice/filter stream fails preflight with a clear message;
  streaming eval is capped (STREAMING_EVAL_MAX_SAMPLES) so each eval terminates;
  the manual-slice shortcut falls back to a regular load when train_split is sliced
- format_conversion: streaming conversions preflight the first mapped row so
  format errors surface before training, not mid-iteration
- frontend: block streaming on Apple Silicon; clear datasetStreaming when a
  dataset is detected as image/audio at start
…eflight test)

- trainer.py: drop unused `IterableDataset` import (hoist safety-net blocker).
- test_training_streaming.py: only select real classes (isinstance type) when
  locating the trainer class, so a MagicMock-stubbed global is never passed to
  object.__new__ (fixes TypeError on the Python 3.10-3.13 jobs).
- no-torch import sandboxes (test_e2e_no_torch_sandbox.py,
  test_studio_import_no_torch.py): teach the chat_templates/format_conversion
  exec stubs and the full-import-chain copy list about the new `.iterable`
  module so the AFTER/runtime cases import without torch again.
@Etherll Etherll closed this Jun 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants