Skip to content

Fix text-only VLM CPT packing truncation - #7211

Merged
danielhanchen merged 16 commits into
unslothai:mainfrom
alkinun:fix/gemma4-text-packing
Jul 20, 2026
Merged

danielhanchen merged 16 commits into
unslothai:mainfrom
alkinun:fix/gemma4-text-packing

Conversation

@alkinun

@alkinun alkinun commented Jul 17, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • allow explicit packing for multimodal-capable models when training uses a raw tokenizer and text-only data
  • keep packing disabled for processors, custom collators, and datasets with vision columns
  • use wrapped packing for Studio CPT and avoid truncating long samples before packing, with compatibility for old and new TRL pack_dataset APIs

Tests

  • python -m pytest tests/utils/test_packing.py -q
  • python -m pytest studio/backend/tests/test_training_preflight.py -q
  • ruff check unsloth/trainer.py unsloth/models/rl_replacements.py studio/backend/core/training/trainer.py tests/utils/test_packing.py

Fixes #7206

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@Etherll Etherll self-assigned this Jul 17, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3c7785eae7

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread unsloth/trainer.py Outdated
@Etherll

Etherll commented Jul 17, 2026

Copy link
Copy Markdown
Collaborator

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b6817905ec

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread unsloth/trainer.py Outdated
Comment thread unsloth/trainer.py Outdated
@Etherll

Etherll commented Jul 17, 2026

Copy link
Copy Markdown
Collaborator

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0d8503eae7

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread unsloth/models/rl_replacements.py Outdated
Comment thread unsloth/trainer.py Outdated
@alkinun

alkinun commented Jul 17, 2026

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cd9a54ee87

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread unsloth/trainer.py
Comment thread unsloth/trainer.py

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bb06e4c8f8

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread unsloth/trainer.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3af93545ff

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread unsloth/models/rl_replacements.py Outdated
Comment thread unsloth/models/rl_replacements.py Outdated
Comment thread unsloth/trainer.py Outdated
@Etherll

Etherll commented Jul 17, 2026

Copy link
Copy Markdown
Collaborator

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. More of your lovely PRs please.

Reviewed commit: cc49802ce4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@danielhanchen

Copy link
Copy Markdown
Member

Thanks for this. Flagging one regression we hit in testing, and I pushed a fix to the branch.

Removing the blanket is_vlm guard also auto-enables padding_free (and packing) for VLM-capable text models whose attention cannot support it. Concretely, unsloth/Qwen3.5-2B is a hybrid linear attention + causal_conv1d model that carries a vision_config, so before this PR it was blocked. With this PR, a text-only run using the raw tokenizer is no longer blocked, so _should_auto_padding_free turns padding_free on. That is the loss difference @Etherll reported.

I checked whether it was just a different but valid loss or real contamination. Using LR=0 (frozen weights, identical batches) on Qwen3.5-2B:

  • batch size 1: padding_free off vs on is bit-identical (it is genuinely active, but a single sequence has nothing to flatten).
  • batch size 2: off vs on diverges up to ~3.6% per step.

The only change between those two cases is a second sequence sharing the padding_free batch, so the linear_attn state and the conv1d kernel leak across the sequence boundary. Same class of problem as the existing gpt_oss entry.

Fix in the pushed commit: add qwen3_5 and qwen3_next to PADDING_FREE_BLOCKLIST, which keeps both padding_free and packing off for these hybrid models. After the change the auto run is blocked and the per-step loss is identical to the padded run again. Standard-attention text VLMs such as Gemma 3 are unaffected and still get the packing fix (verified 100% token retention with finite losses).

Longer term it is probably cleaner to gate the auto-enable on attention type rather than a per-model blocklist, but this keeps the PR correct for now.

@danielhanchen

Copy link
Copy Markdown
Member

Follow-up: generalized the guard so it is not tied to two model names. Replaced the qwen3_5/qwen3_next blocklist entries with a structural check (_is_hybrid_linear_attention_model) that detects any hybrid linear-attention / state-space model: a gated-delta or Mamba-style mixer with a causal conv1d, or a config.layer_types schedule containing linear_attention, or the hybrid config markers.

Verified in an isolated env (transformers 5.2.0, FLA, causal_conv1d): Qwen3.5-2B and its PEFT-wrapped form are detected and have packing/padding_free force-disabled, while Llama-3.2 and Gemma-3 are not flagged and keep packing.

Making packing actually correct for these hybrids (threading seq_idx/cu_seqlens into the causal_conv1d and gated-delta kernels so state resets at packed boundaries) is a larger, transformers-version-dependent change. I will send that as a separate experimental PR behind a flag rather than expand this one.

@danielhanchen

Copy link
Copy Markdown
Member

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 47cc8d4eeb

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread unsloth/trainer.py
Comment on lines +654 to +656
and (
_is_vision_dataset(train_dataset, unknown_is_vision = is_vlm)
or _is_vision_eval_dataset(eval_dataset, unknown_is_vision = is_vlm)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Block packing for images embedded in message columns

When a VLM is passed its raw tokenizer, a dataset using the supported chat format messages=[{"content": [{"type": "image", ...}]}] has no top-level key in _VISION_DATASET_KEYS, so this branch enables packing and padding-free despite containing images. check_dataset_for_missing_videos already recognizes image/video-style message columns in unsloth/models/vision.py; similarly detect nested multimodal content here (or conservatively block messages/conversations) so these runs do not flatten multimodal examples and silently train as text-only.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looked at this closely and it is not reachable as a packing bug. The scenario requires a VLM with an explicitly supplied raw tokenizer (not a processor) plus a messages/conversations dataset carrying nested image content. A PreTrainedTokenizer cannot produce pixel_values, so with a raw tokenizer the images are dropped whether or not packing is on. Blocking packing would not preserve them; the text-only outcome comes from the tokenizer choice, not from flattening. The auto-processor path (processing_class is None) is already blocked via is_auto_processor_vlm, and a processor-supplied VLM is blocked via is_processor. Conservatively blocking messages/conversations would instead disable packing for the text-only chat-format CPT this PR exists to enable (chat-format text is the primary use case), so that would regress the feature. If a defensive warning for nested multimodal content in a raw-tokenizer run is wanted, that is a separate hardening enhancement rather than a correctness fix, and it would need to target nested image/video content specifically, not all chat columns.

Comment thread unsloth/models/rl_replacements.py Outdated
Comment on lines +461 to +463
function = function.replace(
" # All Unsloth Zoo code licensed under LGPLv3\n",
""" # All Unsloth Zoo code licensed under LGPLv3

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Avoid depending on an exact Zoo source header

The matched guard validates only the fast helper's signature, but this injection additionally requires its source to contain this exact license-comment line. With a compatible newer unsloth_zoo that changes or moves that header (the dependency is only lower-bounded), the replacement is a no-op while the later tokenization replacement still introduces _unsloth_wrapped_packing; every affected SFT dataset preparation then fails with NameError. Insert the setup based on a structural location or verify the replacement succeeded before emitting references to these helper variables.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 92f5efe. The _unsloth_wrapped_packing / _inspect setup is now inserted at the sft_prepare_dataset signature (a structural anchor that always exists) via re.subn, instead of matching the exact 'All Unsloth Zoo code licensed under LGPLv3' comment line, and it raises if the signature cannot be located. So the helper variables are always defined before the truncation / pack_dataset rewrites reference them, across unsloth_zoo versions that move or drop that header. Regression test test_wrapped_packing_setup_survives_missing_zoo_header patches in a Zoo source without the header and asserts the setup still lands before the reference.

…omment

The _unsloth_wrapped_packing / _inspect setup block was injected by matching the
exact 'All Unsloth Zoo code licensed under LGPLv3' comment line in the sourced
sft_prepare_dataset. The unsloth_zoo dependency is only lower-bounded, so a newer
Zoo that moves or drops that header made the setup a silent no-op while the
truncation and pack_dataset rewrites still emitted references to those names,
raising NameError on every SFT dataset preparation.

Anchor the setup on the function signature instead (a structural location that
always exists) and fail loudly if it cannot be found, so the helper variables are
always defined before they are referenced across Zoo versions.

Adds a regression test that patches in a Zoo source without the license header.
@danielhanchen

Copy link
Copy Markdown
Member

@codex review

@danielhanchen

Copy link
Copy Markdown
Member

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Nice work!

Reviewed commit: 07fe3bd02e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@danielhanchen
danielhanchen merged commit 9e334d5 into unslothai:main Jul 20, 2026
47 checks passed
danielhanchen added a commit that referenced this pull request Jul 20, 2026
Resolve conflicts after #7211 landed in main. #7249 stacks on #7211 and carries the
more general versions of the shared code, so keep them:
- trainer.py: the varlen shim gating (hybrid_varlen_active), the encoder-decoder
  block, and the string-model config resolution supersede the base hybrid guard.
- rl_replacements.py: the _require_replace helper and _WRAPPED_PACKING_SETUP constant
  supersede the inline function.replace injection.
- test_packing.py: keep the encoder-decoder / decoder-only split and the
  _require_replace / drift-resistant tests; drop the now-duplicated base fixtures.
VectorCipher pushed a commit to VectorCipher/unsloth that referenced this pull request Jul 20, 2026
* Fix text-only VLM CPT packing truncation

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Handle streaming vision datasets in packing

* Harden multimodal packing detection

* Preserve safe packing boundaries

* Scope stream packing checks to VLMs

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Narrow VLM packing detection

* Align packing mode and eval safety

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Add qwen3_5/qwen3_next to PADDING_FREE_BLOCKLIST to avoid packed-sequence contamination

* Detect hybrid linear-attention models structurally instead of by name for packing guard

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Install wrapped-packing setup at the signature, not the Zoo license comment

The _unsloth_wrapped_packing / _inspect setup block was injected by matching the
exact 'All Unsloth Zoo code licensed under LGPLv3' comment line in the sourced
sft_prepare_dataset. The unsloth_zoo dependency is only lower-bounded, so a newer
Zoo that moves or drops that header made the setup a silent no-op while the
truncation and pack_dataset rewrites still emitted references to those names,
raising NameError on every SFT dataset preparation.

Anchor the setup on the function signature instead (a structural location that
always exists) and fail loudly if it cannot be found, so the helper variables are
always defined before they are referenced across Zoo versions.

Adds a regression test that patches in a Zoo source without the license header.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: Etherl <61019402+Etherll@users.noreply.github.com>
Co-authored-by: danielhanchen <danielhanchen@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]Gemma 4 silently disables packing in text-only CPT, causing severe truncation and data loss

3 participants