Repository navigation
chore: prepare v1.0.0 release - #120
Conversation
Summary of ChangesHello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request prepares the project for its v1.0.0 release by performing essential updates across all workspace crates. The primary goal is to transition the project to a stable major version and ensure all necessary metadata is present and correct for publishing to crates.io, enhancing discoverability and compliance. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
- Update all crate versions to 1.0.0 - Add missing metadata to data_connector (repository, keywords) - Add missing metadata to model_gateway (description, license, authors, keywords, categories)
dd1a078 to
5ddd447
Compare
There was a problem hiding this comment.
Code Review
This pull request prepares the workspace for a v1.0.0 release by updating the version of all 14 crates to 1.0.0 and adding necessary package metadata for publishing to crates.io. The changes are straightforward and have been applied correctly across all Cargo.toml files. The version bumps are appropriate for a major release, and the added metadata in data_connector and model_gateway crates will improve their discoverability on crates.io. I've reviewed all the changes and found no issues. The PR is good to merge.
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling now runs against tokenspeed too; model temporarily switched to Qwen/Qwen3-30B-A3B + qwen parser because TokenSpeed's registry does not cover plain LlamaForCausalLM (only LlamaForCausalLMMoE / LlamaForCausalLMEagle3) Test evidence (single B200, Qwen/Qwen3-30B-A3B): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - E2E: tokenspeed 12 passed / 4 failed / 2 skipped vllm (ref) 10 passed / 6 failed / 2 skipped The 4 e2e failures are identical across backends (test_function_call_ required / specific × openai/smg clients). Root cause is upstream of the engine — in smg gateway's tool_choice=required/specific constraint translation path — not this integration. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling now runs against tokenspeed too; model temporarily switched to Qwen/Qwen3-30B-A3B + qwen parser because TokenSpeed's registry does not cover plain LlamaForCausalLM (only LlamaForCausalLMMoE / LlamaForCausalLMEagle3) Test evidence (single B200, Qwen/Qwen3-30B-A3B): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - E2E: tokenspeed 12 passed / 4 failed / 2 skipped vllm (ref) 10 passed / 6 failed / 2 skipped The 4 e2e failures are identical across backends (test_function_call_ required / specific × openai/smg clients). Root cause is upstream of the engine — in smg gateway's tool_choice=required/specific constraint translation path — not this integration. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling keeps its original Llama-3.2-1B + llama parser coverage for sglang/vllm/trtllm; a new TestOpenAIServerFunctionCallingTokenSpeed subclass runs the same test bodies against Qwen/Qwen3-4B + qwen parser for tokenspeed, since TokenSpeed's model registry does not cover plain LlamaForCausalLM (only LlamaForCausalLMMoE / LlamaForCausalLMEagle3). Test evidence (single B200, Qwen/Qwen3-30B-A3B): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - E2E: tokenspeed 12 passed / 4 failed / 2 skipped vllm (ref) 10 passed / 6 failed / 2 skipped The 4 e2e failures are identical across backends (test_function_call_ required / specific × openai/smg clients). Root cause is upstream of the engine — in smg gateway's tool_choice=required/specific constraint translation path — not this integration. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B- Instruct fixture against tokenspeed alongside sglang/vllm/trtllm. TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama`` module registered in the model registry); ci_install_tokenspeed.sh pins to that branch until the upstream PR merges. Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - Registry: ``LlamaForCausalLM`` resolves to the new dense class in the ``tokenspeed:asyncllm`` image; weights load cleanly (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``) - End-to-end generation smoke against the tokenspeed image surfaced an unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm`` build itself (reproduces on stock ``Qwen3-0.6B`` with both the ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix rebuilds TokenSpeed from source, so the relevant end-to-end pass lives in this PR's GitHub Actions run rather than the local image. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B- Instruct fixture against tokenspeed alongside sglang/vllm/trtllm. TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama`` module registered in the model registry); ci_install_tokenspeed.sh pins to that branch until the upstream PR merges. Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - Registry: ``LlamaForCausalLM`` resolves to the new dense class in the ``tokenspeed:asyncllm`` image; weights load cleanly (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``) - End-to-end generation smoke against the tokenspeed image surfaced an unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm`` build itself (reproduces on stock ``Qwen3-0.6B`` with both the ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix rebuilds TokenSpeed from source, so the relevant end-to-end pass lives in this PR's GitHub Actions run rather than the local image. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B- Instruct fixture against tokenspeed alongside sglang/vllm/trtllm. TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama`` module registered in the model registry); ci_install_tokenspeed.sh pins to that branch until the upstream PR merges. Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - Registry: ``LlamaForCausalLM`` resolves to the new dense class in the ``tokenspeed:asyncllm`` image; weights load cleanly (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``) - End-to-end generation smoke against the tokenspeed image surfaced an unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm`` build itself (reproduces on stock ``Qwen3-0.6B`` with both the ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix rebuilds TokenSpeed from source, so the relevant end-to-end pass lives in this PR's GitHub Actions run rather than the local image. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B- Instruct fixture against tokenspeed alongside sglang/vllm/trtllm. TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama`` module registered in the model registry); ci_install_tokenspeed.sh pins to that branch until the upstream PR merges. Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - Registry: ``LlamaForCausalLM`` resolves to the new dense class in the ``tokenspeed:asyncllm`` image; weights load cleanly (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``) - End-to-end generation smoke against the tokenspeed image surfaced an unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm`` build itself (reproduces on stock ``Qwen3-0.6B`` with both the ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix rebuilds TokenSpeed from source, so the relevant end-to-end pass lives in this PR's GitHub Actions run rather than the local image. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing Rust router auto-detects and routes to TokenSpeed workers with zero gateway-side changes. Why the SGLang proto? TokenSpeed and SGLang share the tokenized-request shape, sampling- param naming, and output dict format, so reusing the proto avoids inventing a new one — the previous tokenspeed_scheduler.proto attempt in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this approach. Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/): - servicer.py TokenSpeedSchedulerServicer — Generate, Embed, HealthCheck, Abort, GetModelInfo, GetServerInfo, GetTokenizer, GetLoads - health_servicer.py grpc.health.v1.Health bridge advertising the SGLang service name for router auto-detection - scheduler_launcher.py thin wrapper around TokenSpeed's _launch_subprocesses - server.py python -m entrypoint that boots the scheduler, AsyncLLM, grpc.aio server, and warmup probe - __main__.py CLI shim e2e_test infra: - Runtime.TOKENSPEED enum (infra/constants.py) - _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed workers via python -m smg_grpc_servicer.tokenspeed - Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage - TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B- Instruct fixture against tokenspeed alongside sglang/vllm/trtllm. TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama`` module registered in the model registry); ci_install_tokenspeed.sh pins to that branch until the upstream PR merges. Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture): - Unit: 47 / 47 PASSED (grpc_servicer/tests/) - Registry: ``LlamaForCausalLM`` resolves to the new dense class in the ``tokenspeed:asyncllm`` image; weights load cleanly (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``) - End-to-end generation smoke against the tokenspeed image surfaced an unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm`` build itself (reproduces on stock ``Qwen3-0.6B`` with both the ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix rebuilds TokenSpeed from source, so the relevant end-to-end pass lives in this PR's GitHub Actions run rather than the local image. Closes #120 Signed-off-by: yetone <yetoneful@gmail.com>
Summary
Prepare all workspace crates for v1.0.0 release to crates.io.
Changes
Version Updates
All 14 crates updated from 0.1.0 to 1.0.0:
Metadata Fixes
repositoryandkeywordsfieldsdescription,license,repository,authors,keywords,categoriesfieldsVerification
All crates now have complete metadata:
Test plan
cargo checkpasses