Conversation
for more information, see https://pre-commit.ci
for more information, see https://pre-commit.ci
Reject dataset_streaming at the API boundary when hf_dataset is empty, the dataset is vision/audio, or max_steps is not set. Probe eval split with get_dataset_split_names before the streaming load so typos fail immediately instead of mid-training. Guard column_names=None after map on iterables. Hide the UI toggle for non-text configurations and clear the stale flag when config becomes incompatible.
…emplate/format support (WIP) Work-in-progress on top of feat/studio-dataset-streaming-mode (PR unslothai#4946): - new test_training_streaming.py and iterable.py dataset helper - streaming support in chat_templates.py and format_conversion.py - additional streaming guards in trainer.py / models / routes - frontend streaming wiring in params-section and training-config-store Committed to preserve uncommitted work before merging latest main.
Resolved 4 conflicts from main's CPT/raw-text feature meeting the streaming feature: - trainer.py: keep both imports (detect_streaming_dataset + prepare_raw_text_dataset); keep Optional[int] slice typing and add main's is_cpt param - models/training.py: keep both _validate_dataset_slice and main's field_validators - dataset-section.tsx: drop duplicate barrel imports; keep deep hasSeparateStreamingEvalSplit import - training-config-store.ts: un-fuse two functions' shared tails; merge two version<10 migrations
BLOCKER: streaming + raw-text/CPT crashed on len(IterableDataset). Guard it in the start route (reject format_type=="raw" or training_type=="Continued Pretraining") and in isStreamingSupported (datasetFormat !== "raw"). Also: - models/training.py: validate hf_dataset/subset/split (charset+length, block ..//); cap dataset slice indices (le=1e9); note validator ordering - chat_templates.py: guard _apply_custom_mapping .map() for streaming - trainer.py: warn when packing+streaming - training-config-store.ts: persist-migration bump to v11 (standalone datasetStreaming backfill); add isVisionModel to NON_PERSISTED; toast on silent streamingCompatiblePatch mutations in the 4 indirect setters - tests: route rejections (max_steps, raw/cpt), slice cap, unsafe hf_dataset
Re-merged after main fast-forwarded. Resolved 18 conflicts across 6 files, preserving the streaming feature + review fixes while folding in main's evolution (S3 dataset support, cache-safe load_dataset, base-VLM preflight, manual-slice optimization, improved train-on-responses safety net): - trainer.py (10): combine streaming load path with main's manual-slice fetch; keep IterableDataset import + main's load_dataset_cache_safe; keep both main's _chat_template_renders_empty/_preflight_first_batch and the streaming-aware _train_worker; streaming-skip the post-filter length check; streaming-aware total_steps that also prefers the trainer's processed dataset length - models/training.py: keep dataset_streaming field alongside main's reformat - routes/training.py: keep streaming validation (incl. raw/CPT BLOCKER guard) + main's deliberate-rejection passthrough comment - chat_templates.py / format_conversion.py: keep centralized is_streaming_dataset helper (covers HF + torch) over main's torch-only inline detection - training-config-store.ts: NON_PERSISTED keeps both isVisionModel + s3Config Also fixed a non-conflict auto-merge artifact: duplicate isVisionModel/isAudioModel bindings in dataset-section.tsx. Backend tests: 21/21 pass.
- raw_text: keep the lazy filter but skip len()-based row counting for IterableDatasets so raw-text / CPT can stream; guard the eval-size log - routes/trainer: drop the raw/CPT streaming block; add a defensive not-streaming guard on the eval auto-split (train_test_split) - dataset-section: streaming toggle is visible-but-disabled and lists the exact unmet requirement(s) in its tooltip; block embedding models - training-start-overlay: show "streaming (no full download)" instead of a stuck download bar for streaming runs - trim the streaming test suite to the high-value cases
for more information, see https://pre-commit.ci
…plit, rehydrate timing) - routes: reject dataset_streaming for embedding training and on Apple Silicon (MLX); both loaders materialize the full dataset instead of streaming - trainer: validate the base eval split name so streaming eval accepts HF slice syntax such as "validation[:1000]" - training-config-store: defer the onRehydrateStorage setState to a microtask so it doesn't hit the store's TDZ during synchronous hydration - test: streaming start rejects embedding models
…ty/eval bounds, gating) Address a deeper streaming review: - raw_text: resolve_column_names() guards IterableDataset.column_names=None (from_generator / unresolved features) so raw-text and CPT streaming no longer raise TypeError before training - models/routes: reject HF slice syntax in train_split/eval_split when streaming (load_dataset(streaming=True) raises "Bad split"); reject mixed sources (local/S3) and embedding/MLX streaming at the API, not just in the UI - trainer: an empty post-slice/filter stream fails preflight with a clear message; streaming eval is capped (STREAMING_EVAL_MAX_SAMPLES) so each eval terminates; the manual-slice shortcut falls back to a regular load when train_split is sliced - format_conversion: streaming conversions preflight the first mapped row so format errors surface before training, not mid-iteration - frontend: block streaming on Apple Silicon; clear datasetStreaming when a dataset is detected as image/audio at start
for more information, see https://pre-commit.ci
…eflight test) - trainer.py: drop unused `IterableDataset` import (hoist safety-net blocker). - test_training_streaming.py: only select real classes (isinstance type) when locating the trainer class, so a MagicMock-stubbed global is never passed to object.__new__ (fixes TypeError on the Python 3.10-3.13 jobs). - no-torch import sandboxes (test_e2e_no_torch_sandbox.py, test_studio_import_no_torch.py): teach the chat_templates/format_conversion exec stubs and the full-import-chain copy list about the new `.iterable` module so the AFTER/runtime cases import without torch again.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Staging-only CI validation for unslothai#4946.
This PR exists only to run fork GitHub Actions checks and should be closed after validation.