Repository navigation
refactor: Extract tokenizer into standalone llm-tokenizer workspace crate - #48
Conversation
…rate
Extract the tokenizer module from the main smg crate into a standalone
workspace crate for better modularity, independent versioning, and reusability.
## Changes
### New Crate Structure (tokenizer/)
- Created tokenizer/Cargo.toml with package metadata:
- name: tokenizer (package name)
- lib name: llm_tokenizer (crate name, since "tokenizer" is taken)
- description: LLM tokenizer library with caching and chat template support
- keywords: tokenizer, llm, huggingface, tiktoken, chat-template
- categories: text-processing, parsing
- Moved source files from src/tokenizer/ to tokenizer/src/:
- lib.rs (renamed from mod.rs)
- traits.rs - Core traits (Encoder, Decoder, Tokenizer, Encoding)
- huggingface.rs - HuggingFace tokenizers integration
- tiktoken.rs - OpenAI tiktoken support
- chat_template.rs - Jinja2 chat template processing
- factory.rs - Tokenizer creation utilities
- registry.rs - Thread-safe tokenizer registry
- sequence.rs - Token sequence utilities
- stop.rs - Stop sequence detection
- stream.rs - Streaming decode support
- mock.rs - Mock tokenizer for testing
- hub.rs - HuggingFace Hub download support
- cache/ - Multi-level tokenizer caching:
- mod.rs, fingerprint.rs, l0.rs, l1.rs
### Integration Tests
- Moved tests/tokenizer/ to tokenizer/tests/:
- chat_template_format_detection.rs
- chat_template_integration.rs
- chat_template_loading.rs
- tokenizer_cache_correctness_test.rs
- tokenizer_integration.rs
- Created tests/common/mod.rs with test utilities
### Dependencies
- anyhow, blake3, bytemuck, dashmap, hf-hub
- lru, minijinja (with pycompat), parking_lot, rayon
- serde, serde_json, thiserror, tiktoken-rs, tokenizers
- tokio, tracing, uuid
- Dev: openai-protocol.workspace = true, reqwest, tempfile
### Re-exports from lib.rs
- CacheConfig, CachedTokenizer, CacheStats, L0Cache, L1Cache
- TokenizerFingerprint, TokenizerType, HuggingFaceTokenizer
- MockTokenizer, TokenizerRegistry, LoadError, LoadOutcome
- Sequence, StopSequenceDecoder, StopSequenceConfig
- SequenceDecoderOutput, DecodeStream, TiktokenModel
- Encoder, Decoder, Encoding, SpecialTokens, TokenIdType
### Workspace Configuration
- Added tokenizer to workspace members in root Cargo.toml
- Main crate depends on llm-tokenizer via:
`llm-tokenizer = { path = "tokenizer", package = "tokenizer" }`
- Uses workspace dependency for openai-protocol in dev-dependencies
### Main Crate Changes
- Changed `pub mod tokenizer;` to `pub use llm_tokenizer as tokenizer;`
- Removed src/tokenizer/ directory
- Removed tests/tokenizer/ and tests/tokenizer_tests.rs
## Features
- HuggingFace tokenizers integration
- OpenAI tiktoken support (cl100k, p50k, r50k)
- Jinja2 chat template processing with OpenAI format detection
- Multi-level tokenizer caching (L0 whole-string, L1 prefix)
- Thread-safe tokenizer registry with deduplication
- Stop sequence detection for streaming
- HuggingFace Hub download support
## Notes
- Internal unit tests temporarily disabled pending import fixes
- Integration tests moved but need import updates
Summary of ChangesHello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly refactors the project's tokenizer capabilities by isolating them into a dedicated Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request successfully refactors the tokenizer module into a standalone workspace crate, which is a great step for modularity. The file moves, Cargo.toml updates, and path adjustments seem correct. However, there are two main areas of concern that should be addressed before merging: the removal of the detailed README.md without relocating it to the new crate, and the disabling of the internal unit tests for the new tokenizer crate. Addressing these will ensure the refactoring is complete and maintainable.
I am having trouble creating individual review comments. Click here to see my feedback.
tokenizer/src/lib.rs (23-25)
The internal unit tests for the tokenizer crate have been disabled, as noted in the TODO. A refactoring of this scale should include updates to the corresponding tests to ensure no regressions are introduced. Merging with disabled tests is risky. Please re-enable and fix the tests before this PR is merged.
#[cfg(test)]
mod tests;
src/tokenizer/README.md (1-197)
This detailed README file for the tokenizer module has been removed but not relocated to the new tokenizer crate. This results in a loss of valuable documentation for the new crate. Please move this file to tokenizer/README.md and update its contents (e.g., paths like smg::tokenizer to llm_tokenizer and test commands) to reflect the new crate structure.
1f8fcba to
92c5146
Compare
Re-enabled internal tests that were disabled during crate extraction. Fixed all `use crate::*` glob imports by replacing them with explicit imports, including proper trait imports for Encoder, Decoder, and Tokenizer.
92c5146 to
52936f1
Compare
Extract the tokenizer module from the main smg crate into a standalone workspace crate for better modularity, independent versioning, and reusability.
Changes
New Crate Structure (tokenizer/)
Created tokenizer/Cargo.toml with package metadata:
Moved source files from src/tokenizer/ to tokenizer/src/:
Integration Tests
Dependencies
Re-exports from lib.rs
Workspace Configuration
llm-tokenizer = { path = "tokenizer", package = "tokenizer" }Main Crate Changes
pub mod tokenizer;topub use llm_tokenizer as tokenizer;Features
Notes
fixes #18