Skip to content

Add Hugging Face dataset streaming mode to Studio - #4946

Merged
Etherll merged 24 commits into
unslothai:mainfrom
sanatb187:feat/studio-dataset-streaming-mode
Jun 22, 2026
Merged

Etherll merged 24 commits into
unslothai:mainfrom
sanatb187:feat/studio-dataset-streaming-mode

Conversation

@sanatb187

Copy link
Copy Markdown
Contributor

Summary

Adds Hugging Face dataset streaming mode to Studio.
Implements the Hugging Face dataset streaming part of #4903

Changes

  • Added a Streaming Mode option in the Studio dataset section for Hugging Face datasets
  • Wired the setting through frontend state, API payloads, backend request models, routes, worker, and trainer
  • Added trainer support for streamed HF datasets using iterable loading
  • Added validation for unsupported streaming configurations

Scope

  • This PR supports Hugging Face datasets only
  • Local / recipe-output dataset streaming is left for follow-up work

Validation

  • ruff check . passed locally
  • frontend npm run build passed locally
  • validated on GPU that HF streaming training starts and completes successfully
  • validated explicit eval split works with streaming enabled
  • validated invalid configuration returns an error as expected

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7c648e299c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread studio/backend/core/training/worker.py
Comment thread studio/frontend/src/features/training/stores/training-config-store.ts Outdated
@sanatb187
sanatb187 force-pushed the feat/studio-dataset-streaming-mode branch from 20dc7a6 to 741b1f2 Compare April 10, 2026 07:17

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 640904fecd

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread studio/frontend/src/features/training/stores/training-config-store.ts Outdated
@sanatb187
sanatb187 force-pushed the feat/studio-dataset-streaming-mode branch from 640904f to 7482312 Compare April 10, 2026 08:06

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e1e7218241

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread studio/backend/core/training/trainer.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: aed46bed7b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread studio/backend/core/training/trainer.py

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0139b8c2ad

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread studio/backend/core/training/trainer.py
Roland Tannous and others added 2 commits April 14, 2026 22:37
Reject dataset_streaming at the API boundary when hf_dataset is empty,
the dataset is vision/audio, or max_steps is not set. Probe eval split
with get_dataset_split_names before the streaming load so typos fail
immediately instead of mid-training. Guard column_names=None after map
on iterables. Hide the UI toggle for non-text configurations and clear
the stale flag when config becomes incompatible.
@rolandtannous

rolandtannous commented Apr 14, 2026 •

Copy link
Copy Markdown
Contributor

Pushed 69c85c0a via maintainer edit. Changes:

  • Route validation: reject dataset_streaming=True at the API boundary when hf_dataset is empty, dataset is vision/audio, or max_steps <= 0. Fails before model/dataset loading instead of deep in _train_worker.
  • Eval split probe: get_dataset_split_names call before the streaming eval load so typo'd split names fail immediately with a clear message, not mid-training.
  • .take() warning: IterableDataset.take(N) silently yields fewer rows if the source is shorter — added warning log.
  • Heuristic-rescue guard: IterableDataset.column_names returns None after .map() loses features; guarded the \"text\" in out_ds.column_names check at `dataset_utils.py:1150` against TypeError.
  • UI: streaming toggle hidden for vision/audio models and multimodal datasets; auto-clears `datasetStreaming` when config becomes incompatible.
  • Narrowed _train_worker type hint and updated a couple of error messages.

These changes should be tested extensively before merge to confirm no regression in existing training flows — particularly non-streaming text formats (the heuristic-rescue guard is reachable for any dataset), vision/audio paths (runtime unchanged but UI toggle visibility changed), and the streaming happy path with a real eval split.

@sanatb187

Copy link
Copy Markdown
Contributor Author

Pushed 69c85c0a via maintainer edit. Changes:

  • Route validation: reject dataset_streaming=True at the API boundary when hf_dataset is empty, dataset is vision/audio, or max_steps <= 0. Fails before model/dataset loading instead of deep in _train_worker.
  • Eval split probe: get_dataset_split_names call before the streaming eval load so typo'd split names fail immediately with a clear message, not mid-training.
  • .take() warning: IterableDataset.take(N) silently yields fewer rows if the source is shorter — added warning log.
  • Heuristic-rescue guard: IterableDataset.column_names returns None after .map() loses features; guarded the \"text\" in out_ds.column_names check at dataset_utils.py:1150 against TypeError.
  • UI: streaming toggle hidden for vision/audio models and multimodal datasets; auto-clears datasetStreaming when config becomes incompatible.
  • Narrowed _train_worker type hint and updated a couple of error messages.

These changes should be tested extensively before merge to confirm no regression in existing training flows — particularly non-streaming text formats (the heuristic-rescue guard is reachable for any dataset), vision/audio paths (runtime unchanged but UI toggle visibility changed), and the streaming happy path with a real eval split.

Thanks for the updates @rolandtannous . I see the maintainer edits and that the PR was moved back to draft. I’m reviewing the new changes and will push any needed follow-up fixes before marking it ready again.

@rolandtannous

Copy link
Copy Markdown
Contributor

@sanatb187 there is a more comprehensive fix for sharded models that is going to come first in a separate PR. will tag here. Would appreciate if you could test and note if something is failing.

@Etherll Etherll self-assigned this Jun 15, 2026
Etherll and others added 6 commits June 19, 2026 22:47
…emplate/format support (WIP)

Work-in-progress on top of feat/studio-dataset-streaming-mode (PR unslothai#4946):
- new test_training_streaming.py and iterable.py dataset helper
- streaming support in chat_templates.py and format_conversion.py
- additional streaming guards in trainer.py / models / routes
- frontend streaming wiring in params-section and training-config-store

Committed to preserve uncommitted work before merging latest main.
Resolved 4 conflicts from main's CPT/raw-text feature meeting the streaming feature:
- trainer.py: keep both imports (detect_streaming_dataset + prepare_raw_text_dataset);
  keep Optional[int] slice typing and add main's is_cpt param
- models/training.py: keep both _validate_dataset_slice and main's field_validators
- dataset-section.tsx: drop duplicate barrel imports; keep deep hasSeparateStreamingEvalSplit import
- training-config-store.ts: un-fuse two functions' shared tails; merge two version<10 migrations
BLOCKER: streaming + raw-text/CPT crashed on len(IterableDataset). Guard it in the
start route (reject format_type=="raw" or training_type=="Continued Pretraining")
and in isStreamingSupported (datasetFormat !== "raw").

Also:
- models/training.py: validate hf_dataset/subset/split (charset+length, block ..//);
  cap dataset slice indices (le=1e9); note validator ordering
- chat_templates.py: guard _apply_custom_mapping .map() for streaming
- trainer.py: warn when packing+streaming
- training-config-store.ts: persist-migration bump to v11 (standalone datasetStreaming
  backfill); add isVisionModel to NON_PERSISTED; toast on silent streamingCompatiblePatch
  mutations in the 4 indirect setters
- tests: route rejections (max_steps, raw/cpt), slice cap, unsafe hf_dataset
Re-merged after main fast-forwarded. Resolved 18 conflicts across 6 files,
preserving the streaming feature + review fixes while folding in main's evolution
(S3 dataset support, cache-safe load_dataset, base-VLM preflight, manual-slice
optimization, improved train-on-responses safety net):

- trainer.py (10): combine streaming load path with main's manual-slice fetch;
  keep IterableDataset import + main's load_dataset_cache_safe; keep both main's
  _chat_template_renders_empty/_preflight_first_batch and the streaming-aware
  _train_worker; streaming-skip the post-filter length check; streaming-aware
  total_steps that also prefers the trainer's processed dataset length
- models/training.py: keep dataset_streaming field alongside main's reformat
- routes/training.py: keep streaming validation (incl. raw/CPT BLOCKER guard) +
  main's deliberate-rejection passthrough comment
- chat_templates.py / format_conversion.py: keep centralized is_streaming_dataset
  helper (covers HF + torch) over main's torch-only inline detection
- training-config-store.ts: NON_PERSISTED keeps both isVisionModel + s3Config

Also fixed a non-conflict auto-merge artifact: duplicate isVisionModel/isAudioModel
bindings in dataset-section.tsx. Backend tests: 21/21 pass.
- raw_text: keep the lazy filter but skip len()-based row counting for
  IterableDatasets so raw-text / CPT can stream; guard the eval-size log
- routes/trainer: drop the raw/CPT streaming block; add a defensive
  not-streaming guard on the eval auto-split (train_test_split)
- dataset-section: streaming toggle is visible-but-disabled and lists the
  exact unmet requirement(s) in its tooltip; block embedding models
- training-start-overlay: show "streaming (no full download)" instead of a
  stuck download bar for streaming runs
- trim the streaming test suite to the high-value cases
@Etherll
Etherll marked this pull request as ready for review June 19, 2026 19:54
@Etherll
Etherll requested a review from danielhanchen as a code owner June 19, 2026 19:54

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a58949279a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread studio/backend/routes/training.py
Comment thread studio/backend/routes/training.py
Comment thread studio/backend/core/training/trainer.py
Comment thread studio/frontend/src/features/training/stores/training-config-store.ts Outdated
Etherll and others added 3 commits June 19, 2026 23:18
…plit, rehydrate timing)

- routes: reject dataset_streaming for embedding training and on Apple Silicon
  (MLX); both loaders materialize the full dataset instead of streaming
- trainer: validate the base eval split name so streaming eval accepts HF slice
  syntax such as "validation[:1000]"
- training-config-store: defer the onRehydrateStorage setState to a microtask so
  it doesn't hit the store's TDZ during synchronous hydration
- test: streaming start rejects embedding models
…ty/eval bounds, gating)

Address a deeper streaming review:
- raw_text: resolve_column_names() guards IterableDataset.column_names=None
  (from_generator / unresolved features) so raw-text and CPT streaming no longer
  raise TypeError before training
- models/routes: reject HF slice syntax in train_split/eval_split when streaming
  (load_dataset(streaming=True) raises "Bad split"); reject mixed sources
  (local/S3) and embedding/MLX streaming at the API, not just in the UI
- trainer: an empty post-slice/filter stream fails preflight with a clear message;
  streaming eval is capped (STREAMING_EVAL_MAX_SAMPLES) so each eval terminates;
  the manual-slice shortcut falls back to a regular load when train_split is sliced
- format_conversion: streaming conversions preflight the first mapped row so
  format errors surface before training, not mid-iteration
- frontend: block streaming on Apple Silicon; clear datasetStreaming when a
  dataset is detected as image/audio at start

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 14b7d30bc3

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread studio/frontend/src/features/training/hooks/use-training-actions.ts
@Etherll

Etherll commented Jun 20, 2026

Copy link
Copy Markdown
Collaborator

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. More of your lovely PRs please.

Reviewed commit: 8820ef1d4e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

…eflight test)

- trainer.py: drop unused `IterableDataset` import (hoist safety-net blocker).
- test_training_streaming.py: only select real classes (isinstance type) when
  locating the trainer class, so a MagicMock-stubbed global is never passed to
  object.__new__ (fixes TypeError on the Python 3.10-3.13 jobs).
- no-torch import sandboxes (test_e2e_no_torch_sandbox.py,
  test_studio_import_no_torch.py): teach the chat_templates/format_conversion
  exec stubs and the full-import-chain copy list about the new `.iterable`
  module so the AFTER/runtime cases import without torch again.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@sanatb187

Copy link
Copy Markdown
Contributor Author

Thank you for the support on this @Etherll !

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@Etherll

Etherll commented Jun 22, 2026

Copy link
Copy Markdown
Collaborator

Thank you for the support on this @Etherll !

Thanks alot of the PR good job !

@Etherll Etherll left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM !

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.

@Etherll
Etherll merged commit 1fc8bf5 into unslothai:main Jun 22, 2026
1 check was pending
@sanatb187
sanatb187 deleted the feat/studio-dataset-streaming-mode branch June 22, 2026 14:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing WIP We're working on it

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants