Skip to content

refactor: Extract tokenizer into standalone llm-tokenizer workspace crate - #48

Merged
slin1237 merged 2 commits into
mainfrom
tokenizer-ex
Jan 19, 2026
Merged

slin1237 merged 2 commits into
mainfrom
tokenizer-ex

Conversation

@slin1237

@slin1237 slin1237 commented Jan 19, 2026 •

Copy link
Copy Markdown
Member

Extract the tokenizer module from the main smg crate into a standalone workspace crate for better modularity, independent versioning, and reusability.

Changes

New Crate Structure (tokenizer/)

  • Created tokenizer/Cargo.toml with package metadata:

    • name: tokenizer (package name)
    • lib name: llm_tokenizer (crate name, since "tokenizer" is taken)
    • description: LLM tokenizer library with caching and chat template support
    • keywords: tokenizer, llm, huggingface, tiktoken, chat-template
    • categories: text-processing, parsing
  • Moved source files from src/tokenizer/ to tokenizer/src/:

    • lib.rs (renamed from mod.rs)
    • traits.rs - Core traits (Encoder, Decoder, Tokenizer, Encoding)
    • huggingface.rs - HuggingFace tokenizers integration
    • tiktoken.rs - OpenAI tiktoken support
    • chat_template.rs - Jinja2 chat template processing
    • factory.rs - Tokenizer creation utilities
    • registry.rs - Thread-safe tokenizer registry
    • sequence.rs - Token sequence utilities
    • stop.rs - Stop sequence detection
    • stream.rs - Streaming decode support
    • mock.rs - Mock tokenizer for testing
    • hub.rs - HuggingFace Hub download support
    • cache/ - Multi-level tokenizer caching:
      • mod.rs, fingerprint.rs, l0.rs, l1.rs

Integration Tests

  • Moved tests/tokenizer/ to tokenizer/tests/:
    • chat_template_format_detection.rs
    • chat_template_integration.rs
    • chat_template_loading.rs
    • tokenizer_cache_correctness_test.rs
    • tokenizer_integration.rs
    • Created tests/common/mod.rs with test utilities

Dependencies

  • anyhow, blake3, bytemuck, dashmap, hf-hub
  • lru, minijinja (with pycompat), parking_lot, rayon
  • serde, serde_json, thiserror, tiktoken-rs, tokenizers
  • tokio, tracing, uuid
  • Dev: openai-protocol.workspace = true, reqwest, tempfile

Re-exports from lib.rs

  • CacheConfig, CachedTokenizer, CacheStats, L0Cache, L1Cache
  • TokenizerFingerprint, TokenizerType, HuggingFaceTokenizer
  • MockTokenizer, TokenizerRegistry, LoadError, LoadOutcome
  • Sequence, StopSequenceDecoder, StopSequenceConfig
  • SequenceDecoderOutput, DecodeStream, TiktokenModel
  • Encoder, Decoder, Encoding, SpecialTokens, TokenIdType

Workspace Configuration

  • Added tokenizer to workspace members in root Cargo.toml
  • Main crate depends on llm-tokenizer via: llm-tokenizer = { path = "tokenizer", package = "tokenizer" }
  • Uses workspace dependency for openai-protocol in dev-dependencies

Main Crate Changes

  • Changed pub mod tokenizer; to pub use llm_tokenizer as tokenizer;
  • Removed src/tokenizer/ directory
  • Removed tests/tokenizer/ and tests/tokenizer_tests.rs

Features

  • HuggingFace tokenizers integration
  • OpenAI tiktoken support (cl100k, p50k, r50k)
  • Jinja2 chat template processing with OpenAI format detection
  • Multi-level tokenizer caching (L0 whole-string, L1 prefix)
  • Thread-safe tokenizer registry with deduplication
  • Stop sequence detection for streaming
  • HuggingFace Hub download support

Notes

  • Internal unit tests temporarily disabled pending import fixes
  • Integration tests moved but need import updates

fixes #18

…rate

Extract the tokenizer module from the main smg crate into a standalone
workspace crate for better modularity, independent versioning, and reusability.

## Changes

### New Crate Structure (tokenizer/)
- Created tokenizer/Cargo.toml with package metadata:
  - name: tokenizer (package name)
  - lib name: llm_tokenizer (crate name, since "tokenizer" is taken)
  - description: LLM tokenizer library with caching and chat template support
  - keywords: tokenizer, llm, huggingface, tiktoken, chat-template
  - categories: text-processing, parsing

- Moved source files from src/tokenizer/ to tokenizer/src/:
  - lib.rs (renamed from mod.rs)
  - traits.rs - Core traits (Encoder, Decoder, Tokenizer, Encoding)
  - huggingface.rs - HuggingFace tokenizers integration
  - tiktoken.rs - OpenAI tiktoken support
  - chat_template.rs - Jinja2 chat template processing
  - factory.rs - Tokenizer creation utilities
  - registry.rs - Thread-safe tokenizer registry
  - sequence.rs - Token sequence utilities
  - stop.rs - Stop sequence detection
  - stream.rs - Streaming decode support
  - mock.rs - Mock tokenizer for testing
  - hub.rs - HuggingFace Hub download support
  - cache/ - Multi-level tokenizer caching:
    - mod.rs, fingerprint.rs, l0.rs, l1.rs

### Integration Tests
- Moved tests/tokenizer/ to tokenizer/tests/:
  - chat_template_format_detection.rs
  - chat_template_integration.rs
  - chat_template_loading.rs
  - tokenizer_cache_correctness_test.rs
  - tokenizer_integration.rs
  - Created tests/common/mod.rs with test utilities

### Dependencies
- anyhow, blake3, bytemuck, dashmap, hf-hub
- lru, minijinja (with pycompat), parking_lot, rayon
- serde, serde_json, thiserror, tiktoken-rs, tokenizers
- tokio, tracing, uuid
- Dev: openai-protocol.workspace = true, reqwest, tempfile

### Re-exports from lib.rs
- CacheConfig, CachedTokenizer, CacheStats, L0Cache, L1Cache
- TokenizerFingerprint, TokenizerType, HuggingFaceTokenizer
- MockTokenizer, TokenizerRegistry, LoadError, LoadOutcome
- Sequence, StopSequenceDecoder, StopSequenceConfig
- SequenceDecoderOutput, DecodeStream, TiktokenModel
- Encoder, Decoder, Encoding, SpecialTokens, TokenIdType

### Workspace Configuration
- Added tokenizer to workspace members in root Cargo.toml
- Main crate depends on llm-tokenizer via:
  `llm-tokenizer = { path = "tokenizer", package = "tokenizer" }`
- Uses workspace dependency for openai-protocol in dev-dependencies

### Main Crate Changes
- Changed `pub mod tokenizer;` to `pub use llm_tokenizer as tokenizer;`
- Removed src/tokenizer/ directory
- Removed tests/tokenizer/ and tests/tokenizer_tests.rs

## Features
- HuggingFace tokenizers integration
- OpenAI tiktoken support (cl100k, p50k, r50k)
- Jinja2 chat template processing with OpenAI format detection
- Multi-level tokenizer caching (L0 whole-string, L1 prefix)
- Thread-safe tokenizer registry with deduplication
- Stop sequence detection for streaming
- HuggingFace Hub download support

## Notes
- Internal unit tests temporarily disabled pending import fixes
- Integration tests moved but need import updates
@github-actions github-actions Bot added documentation Improvements or additions to documentation tokenizer Tokenizer related changes rust dependencies Dependency updates tests Test changes labels Jan 19, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly refactors the project's tokenizer capabilities by isolating them into a dedicated llm-tokenizer crate within the workspace. This strategic move aims to improve the overall architecture by promoting better separation of concerns, allowing the tokenizer component to evolve independently, and making it easier to reuse this crucial functionality across different parts of the system or in other projects. The core logic for tokenization, chat template handling, and caching remains consistent but is now encapsulated in its own well-defined boundary.

Highlights

  • Module Extraction: The tokenizer module has been extracted from the main smg crate into a new, standalone workspace crate named llm-tokenizer. This enhances modularity, independent versioning, and reusability.
  • New Crate Structure: A new directory tokenizer/ has been created, containing its own Cargo.toml and all tokenizer-related source files, now located under tokenizer/src/. The library name for the new crate is llm_tokenizer.
  • Dependency and Import Updates: The root Cargo.toml has been updated to include llm-tokenizer as a workspace member and a dependency for the main crate. All internal use statements within the tokenizer files have been adjusted to reflect the new crate path, and the main crate now re-exports llm_tokenizer as tokenizer.
  • Test Relocation and Adjustments: All tokenizer integration tests have been moved from tests/tokenizer/ to tokenizer/tests/. A new common test utility file (tokenizer/tests/common/mod.rs) was added. Note that internal unit tests are temporarily disabled, and integration tests require further import updates.
  • Exposed Features: The new llm-tokenizer crate exposes comprehensive features including HuggingFace and OpenAI Tiktoken support, Jinja2 chat template processing, multi-level caching, a thread-safe tokenizer registry, and stop sequence detection for streaming.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@slin1237 slin1237 added the run-ci label to trigger ci workflow label Jan 19, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request successfully refactors the tokenizer module into a standalone workspace crate, which is a great step for modularity. The file moves, Cargo.toml updates, and path adjustments seem correct. However, there are two main areas of concern that should be addressed before merging: the removal of the detailed README.md without relocating it to the new crate, and the disabling of the internal unit tests for the new tokenizer crate. Addressing these will ensure the refactoring is complete and maintainable.

I am having trouble creating individual review comments. Click here to see my feedback.

tokenizer/src/lib.rs (23-25)

high

The internal unit tests for the tokenizer crate have been disabled, as noted in the TODO. A refactoring of this scale should include updates to the corresponding tests to ensure no regressions are introduced. Merging with disabled tests is risky. Please re-enable and fix the tests before this PR is merged.

#[cfg(test)]
mod tests;

src/tokenizer/README.md (1-197)

medium

This detailed README file for the tokenizer module has been removed but not relocated to the new tokenizer crate. This results in a loss of valuable documentation for the new crate. Please move this file to tokenizer/README.md and update its contents (e.g., paths like smg::tokenizer to llm_tokenizer and test commands) to reflect the new crate structure.

Re-enabled internal tests that were disabled during crate extraction.
Fixed all `use crate::*` glob imports by replacing them with explicit
imports, including proper trait imports for Encoder, Decoder, and
Tokenizer.
@github-actions github-actions Bot removed the tokenizer Tokenizer related changes label Jan 19, 2026
@slin1237
slin1237 merged commit aa1efb1 into main Jan 19, 2026
11 of 12 checks passed
@slin1237
slin1237 deleted the tokenizer-ex branch January 19, 2026 05:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Dependency updates documentation Improvements or additions to documentation run-ci label to trigger ci workflow tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Extract tokenizer crate

1 participant