Skip to content

Add VarlenLowLevelDataset, VarlenDataset, and MockVarlenDataset - #5909

Closed
ilml wants to merge 2 commits into
NVIDIA:pull-request/5907from
ilml:split/3386-09-varlen-datasets
Closed

Add VarlenLowLevelDataset, VarlenDataset, and MockVarlenDataset#5909
ilml wants to merge 2 commits into
NVIDIA:pull-request/5907from
ilml:split/3386-09-varlen-datasets

Conversation

@ilml

@ilml ilml commented Jul 20, 2026

Copy link
Copy Markdown
Contributor
  • I, the PR author, have personally reviewed every line of this PR.

What does this PR do?

Adds the variable-length packed (THD) dataset family: HF-hub/parquet/jsonl loading, THD __getitem__ emitting cu_seqlens/original_seq_len/padded_seq_len, the SBHD validation mode, and the mock variant, with unit tests.

Part 09/10 of splitting #3386 (Add E2E support for THD format; dev-branch PR #2924). Original changes by @xiaoyao0115 in #3386 — split into functionally self-contained PRs to ease review. Hard dependencies (must merge first): #5906, #5908.

Split series (#3386)

Part PR Hard deps (merge first)
01 #5901
02 #5902
03 #5903
04 #6589 #5903
05 #5905 #6589
06 #5906 #5902, #5903
07 #5907 #5902, #6589, #5905
08 #5908
09 #5909 #5906, #5908
10 #5910 #5907, #5909

Branches are stacked linearly (each on the previous) so every PR shows a clean own-diff once its base is retargeted to the copy-pr-bot pull-request/<parent> ref; until then the Files-changed view of a stacked PR includes its ancestors — its own change is the last commit.
Stacked on #5907 (carries a copy of #5908's commit so the branch is self-sufficient; it drops out automatically when #5908 merges).

Review notes

  • Deviation from Add E2E support for THD format #3386: adds hybrid_context_parallel=False to the _make_config test helper — without it, six __getitem__ tests crash with AttributeError because _calculate_padding_divisor reads config.hybrid_context_parallel unguarded on a SimpleNamespace config.

Issue tracking

Linked issue: Related to #3386

Contribution process

Pre-checks

  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

🤖 Generated with Claude Code

@copy-pr-bot

copy-pr-bot Bot commented Jul 20, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

ilml and others added 2 commits August 10, 2026 22:09
Split 8/10 from NVIDIA#3386 (sequence packing / THD E2E support). Adds the
pure-Python schema detection (openai-messages, sharegpt, alpaca/dolly,
pretraining-text) and message-normalization helpers for the varlen
dataset, with CPU-only unit tests.

Original changes by @xiaoyao0115 in NVIDIA#3386.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: ilml <tolong@nvidia.com>
Split 9/10 from NVIDIA#3386 (sequence packing / THD E2E support). Adds the
variable-length packed (THD) dataset family: HF-hub/parquet/jsonl
loading, THD __getitem__ with cu_seqlens, SBHD validation mode, and the
mock variant, with unit tests.

Also adds hybrid_context_parallel=False to the _make_config test helper
(deviation from NVIDIA#3386: fixes a latent AttributeError in
_calculate_padding_divisor with SimpleNamespace configs).

Original changes by @xiaoyao0115 in NVIDIA#3386.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: ilml <tolong@nvidia.com>
@ilml

ilml commented Aug 17, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test 28eed6b

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants