fix: guard remove_special_tokens against tokenizers without a BOS token - #7048
Merged
oobabooga merged 2 commits intoJul 10, 2026
Merged
Conversation
remove_special_tokens called prompt.startswith(tokenizer.bos_token) without checking bos_token first. Tokenizers such as Qwen2/Qwen2.5, GPT-2, Falcon and GPT-NeoX have no BOS token (tokenizer.bos_token is None), so the call raised "TypeError: startswith first arg must be str or a tuple of str, not NoneType" and crashed test_hf_gguf_equivalence for those models. Guard bos_token with getattr(tokenizer, "bos_token", None), mirroring the None checks already used elsewhere in this file (get_ollama_eos_tokens and _change_system_message).
for more information, see https://pre-commit.ci
Contributor
|
Note Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported. |
Member
|
Confirmed this on Thanks @vineethsaivs. Merging. |
This was referenced Jul 22, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
remove_special_tokensstrips a leading BOS token but callsprompt.startswith(tokenizer.bos_token)without checking that the tokenizer actually has a BOS token. Many popular tokenizers, including Qwen2 / Qwen2.5, GPT-2, Falcon and GPT-NeoX, define no BOS token, sotokenizer.bos_tokenisNoneand the call raises:remove_special_tokensis the final step oftest_hf_gguf_equivalence, so running the HF vs GGUF equivalence check on any of those models crashes.Fix
Guard
bos_tokenwithgetattr(tokenizer, "bos_token", None)before the comparison, mirroring theNonechecks already used elsewhere in this same file (get_ollama_eos_tokensand_change_system_message). When the tokenizer has no BOS token the prompt is returned unchanged; the existing single-BOS-stripping behaviour is preserved for tokenizers that do have one.Test
tests/python/test_remove_special_tokens_no_bos.pyloads the function viaast(no GPU, noimport unsloth) and covers:bos_token=None) returns the prompt unchanged (this raises theTypeErrorbefore the fix),