Skip to content

chore: prepare v1.0.0 release - #120

Merged
slin1237 merged 1 commit into
mainfrom
release/v1.0.0
Jan 21, 2026
Merged

slin1237 merged 1 commit into
mainfrom
release/v1.0.0

Conversation

@slin1237

Copy link
Copy Markdown
Member

Summary

Prepare all workspace crates for v1.0.0 release to crates.io.

Changes

Version Updates

All 14 crates updated from 0.1.0 to 1.0.0:

  • openai-protocol
  • reasoning-parser
  • tool-parser
  • wfaas
  • llm-tokenizer
  • auth
  • smg-mcp
  • kv-index
  • data-connector
  • llm-multimodal
  • smg-wasm
  • smg-mesh
  • smg-grpc-client
  • smg (main gateway)

Metadata Fixes

  • data_connector: Added repository and keywords fields
  • model_gateway: Added description, license, repository, authors, keywords, categories fields

Verification

All crates now have complete metadata:

  • ✅ Authors
  • ✅ License (Apache-2.0)
  • ✅ Keywords
  • ✅ Repository URL
  • ✅ Description
  • ✅ Categories

Test plan

  • cargo check passes
  • Merge and verify crate publishing workflows trigger

@github-actions github-actions Bot added tokenizer Tokenizer related changes dependencies Dependency updates grpc gRPC client and router changes mcp MCP related changes wasm WebAssembly related changes tool-parser Tool/function call parser changes reasoning-parser Reasoning parser changes auth Authentication crate changes data-connector Data connector crate changes multimodal Multimodal crate changes protocols Protocols crate changes workflow Workflow crate changes model-gateway Model gateway crate changes labels Jan 21, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request prepares the project for its v1.0.0 release by performing essential updates across all workspace crates. The primary goal is to transition the project to a stable major version and ensure all necessary metadata is present and correct for publishing to crates.io, enhancing discoverability and compliance.

Highlights

  • Version Updates: All 14 workspace crates have been updated to version "1.0.0" from their previous "0.1.0" or "0.3.2" versions, signifying a stable release.
  • Metadata Enhancements: The "data_connector" crate now includes "repository" and "keywords" fields, while the "model_gateway" crate has been enriched with "description", "license", "repository", "authors", "keywords", and "categories" to meet crates.io publishing requirements.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

- Update all crate versions to 1.0.0
- Add missing metadata to data_connector (repository, keywords)
- Add missing metadata to model_gateway (description, license, authors, keywords, categories)

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request prepares the workspace for a v1.0.0 release by updating the version of all 14 crates to 1.0.0 and adding necessary package metadata for publishing to crates.io. The changes are straightforward and have been applied correctly across all Cargo.toml files. The version bumps are appropriate for a major release, and the added metadata in data_connector and model_gateway crates will improve their discoverability on crates.io. I've reviewed all the changes and found no issues. The PR is good to merge.

@slin1237
slin1237 merged commit e848296 into main Jan 21, 2026
16 checks passed
@slin1237
slin1237 deleted the release/v1.0.0 branch January 21, 2026 02:08
yetone added a commit that referenced this pull request Apr 23, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling now runs against tokenspeed too;
  model temporarily switched to Qwen/Qwen3-30B-A3B + qwen parser
  because TokenSpeed's registry does not cover plain
  LlamaForCausalLM (only LlamaForCausalLMMoE / LlamaForCausalLMEagle3)

Test evidence (single B200, Qwen/Qwen3-30B-A3B):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- E2E:   tokenspeed 12 passed / 4 failed / 2 skipped
         vllm (ref) 10 passed / 6 failed / 2 skipped

The 4 e2e failures are identical across backends (test_function_call_
required / specific × openai/smg clients). Root cause is upstream of
the engine — in smg gateway's tool_choice=required/specific
constraint translation path — not this integration.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
yetone added a commit that referenced this pull request Apr 23, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling now runs against tokenspeed too;
  model temporarily switched to Qwen/Qwen3-30B-A3B + qwen parser
  because TokenSpeed's registry does not cover plain
  LlamaForCausalLM (only LlamaForCausalLMMoE / LlamaForCausalLMEagle3)

Test evidence (single B200, Qwen/Qwen3-30B-A3B):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- E2E:   tokenspeed 12 passed / 4 failed / 2 skipped
         vllm (ref) 10 passed / 6 failed / 2 skipped

The 4 e2e failures are identical across backends (test_function_call_
required / specific × openai/smg clients). Root cause is upstream of
the engine — in smg gateway's tool_choice=required/specific
constraint translation path — not this integration.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
yetone added a commit that referenced this pull request Apr 23, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling keeps its original Llama-3.2-1B +
  llama parser coverage for sglang/vllm/trtllm; a new
  TestOpenAIServerFunctionCallingTokenSpeed subclass runs the same
  test bodies against Qwen/Qwen3-4B + qwen parser for tokenspeed,
  since TokenSpeed's model registry does not cover plain
  LlamaForCausalLM (only LlamaForCausalLMMoE / LlamaForCausalLMEagle3).

Test evidence (single B200, Qwen/Qwen3-30B-A3B):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- E2E:   tokenspeed 12 passed / 4 failed / 2 skipped
         vllm (ref) 10 passed / 6 failed / 2 skipped

The 4 e2e failures are identical across backends (test_function_call_
required / specific × openai/smg clients). Root cause is upstream of
the engine — in smg gateway's tool_choice=required/specific
constraint translation path — not this integration.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
yetone added a commit that referenced this pull request Apr 23, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B-
  Instruct fixture against tokenspeed alongside sglang/vllm/trtllm.
  TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via
  lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama``
  module registered in the model registry); ci_install_tokenspeed.sh
  pins to that branch until the upstream PR merges.

Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- Registry: ``LlamaForCausalLM`` resolves to the new dense class in the
  ``tokenspeed:asyncllm`` image; weights load cleanly
  (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``)
- End-to-end generation smoke against the tokenspeed image surfaced an
  unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm``
  build itself (reproduces on stock ``Qwen3-0.6B`` with both the
  ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix
  rebuilds TokenSpeed from source, so the relevant end-to-end pass
  lives in this PR's GitHub Actions run rather than the local image.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
yetone added a commit that referenced this pull request Apr 24, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B-
  Instruct fixture against tokenspeed alongside sglang/vllm/trtllm.
  TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via
  lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama``
  module registered in the model registry); ci_install_tokenspeed.sh
  pins to that branch until the upstream PR merges.

Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- Registry: ``LlamaForCausalLM`` resolves to the new dense class in the
  ``tokenspeed:asyncllm`` image; weights load cleanly
  (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``)
- End-to-end generation smoke against the tokenspeed image surfaced an
  unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm``
  build itself (reproduces on stock ``Qwen3-0.6B`` with both the
  ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix
  rebuilds TokenSpeed from source, so the relevant end-to-end pass
  lives in this PR's GitHub Actions run rather than the local image.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
yetone added a commit that referenced this pull request Apr 27, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B-
  Instruct fixture against tokenspeed alongside sglang/vllm/trtllm.
  TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via
  lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama``
  module registered in the model registry); ci_install_tokenspeed.sh
  pins to that branch until the upstream PR merges.

Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- Registry: ``LlamaForCausalLM`` resolves to the new dense class in the
  ``tokenspeed:asyncllm`` image; weights load cleanly
  (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``)
- End-to-end generation smoke against the tokenspeed image surfaced an
  unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm``
  build itself (reproduces on stock ``Qwen3-0.6B`` with both the
  ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix
  rebuilds TokenSpeed from source, so the relevant end-to-end pass
  lives in this PR's GitHub Actions run rather than the local image.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
CatherineSue pushed a commit that referenced this pull request Apr 30, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B-
  Instruct fixture against tokenspeed alongside sglang/vllm/trtllm.
  TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via
  lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama``
  module registered in the model registry); ci_install_tokenspeed.sh
  pins to that branch until the upstream PR merges.

Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- Registry: ``LlamaForCausalLM`` resolves to the new dense class in the
  ``tokenspeed:asyncllm`` image; weights load cleanly
  (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``)
- End-to-end generation smoke against the tokenspeed image surfaced an
  unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm``
  build itself (reproduces on stock ``Qwen3-0.6B`` with both the
  ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix
  rebuilds TokenSpeed from source, so the relevant end-to-end pass
  lives in this PR's GitHub Actions run rather than the local image.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
key4ng pushed a commit that referenced this pull request Apr 30, 2026
Wraps TokenSpeed's AsyncLLM behind the SGLang proto so the existing
Rust router auto-detects and routes to TokenSpeed workers with zero
gateway-side changes.

Why the SGLang proto?
TokenSpeed and SGLang share the tokenized-request shape, sampling-
param naming, and output dict format, so reusing the proto avoids
inventing a new one — the previous tokenspeed_scheduler.proto attempt
in lightseekorg/tokenspeed#185 and #1167 was closed in favor of this
approach.

Module layout (grpc_servicer/smg_grpc_servicer/tokenspeed/):
- servicer.py           TokenSpeedSchedulerServicer — Generate, Embed,
                        HealthCheck, Abort, GetModelInfo,
                        GetServerInfo, GetTokenizer, GetLoads
- health_servicer.py    grpc.health.v1.Health bridge advertising the
                        SGLang service name for router auto-detection
- scheduler_launcher.py thin wrapper around TokenSpeed's
                        _launch_subprocesses
- server.py             python -m entrypoint that boots the scheduler,
                        AsyncLLM, grpc.aio server, and warmup probe
- __main__.py           CLI shim

e2e_test infra:
- Runtime.TOKENSPEED enum (infra/constants.py)
- _build_tokenspeed_grpc_cmd so e2e fixtures can spawn tokenspeed
  workers via python -m smg_grpc_servicer.tokenspeed
- Qwen/Qwen3-4B and Qwen/Qwen3-30B-A3B model specs for coverage
- TestOpenAIServerFunctionCalling now runs the shared Llama-3.2-1B-
  Instruct fixture against tokenspeed alongside sglang/vllm/trtllm.
  TokenSpeed support for dense ``LlamaForCausalLM`` is wired up via
  lightseekorg/tokenspeed#357 (a new ``tokenspeed.runtime.models.llama``
  module registered in the model registry); ci_install_tokenspeed.sh
  pins to that branch until the upstream PR merges.

Test evidence (single B200, shared Llama-3.2-1B-Instruct fixture):
- Unit:  47 / 47 PASSED (grpc_servicer/tests/)
- Registry: ``LlamaForCausalLM`` resolves to the new dense class in the
  ``tokenspeed:asyncllm`` image; weights load cleanly
  (``Load weight end. type=LlamaForCausalLM, dtype=torch.bfloat16``)
- End-to-end generation smoke against the tokenspeed image surfaced an
  unrelated detokenizer/sampler issue in the ``tokenspeed:asyncllm``
  build itself (reproduces on stock ``Qwen3-0.6B`` with both the
  ``greedy`` and ``flashinfer`` sampling backends). The CI e2e matrix
  rebuilds TokenSpeed from source, so the relevant end-to-end pass
  lives in this PR's GitHub Actions run rather than the local image.

Closes #120

Signed-off-by: yetone <yetoneful@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auth Authentication crate changes data-connector Data connector crate changes dependencies Dependency updates grpc gRPC client and router changes mcp MCP related changes model-gateway Model gateway crate changes multimodal Multimodal crate changes protocols Protocols crate changes reasoning-parser Reasoning parser changes tokenizer Tokenizer related changes tool-parser Tool/function call parser changes wasm WebAssembly related changes workflow Workflow crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant