Skip to content

[Doc][Misc] Align Kimi-K3 tutorial with the deployment template - #15186

Merged
liqian79 merged 3 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-docs
Aug 28, 2026
Merged

liqian79 merged 3 commits into
vllm-project:mainfrom
maoxx241:codex/kimi-k3-docs

Conversation

@maoxx241

@maoxx241 maoxx241 commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Follow up on #14454 to align the Kimi-K3 documentation with the model deployment tutorial template and the existing Kimi, DeepSeek, and Qwen tutorials.

  • Add Kimi-K3 to the MkDocs model navigation and bilingual title mapping, and restore the generic model tutorial overview.
  • Add a conservative A3 W4A8 entry to the supported-model matrix and link the tutorial to the shared model/feature guides.
  • Organize the tutorial into the standard ten sections, including prerequisites, installation, full-checkpoint deployment, functional verification, accuracy/performance evaluation, tuning, and FAQ.
  • Document the Atlas A3-only installation scope, main-branch A3 image, and source-build path using the verified upstream vLLM revision; add key serving-parameter explanations and a reference configuration table.
  • Provide ModelScope links for the Eco-Tech W4A8 target, RadixArk GQA draft, and both Inferact MLA drafts. Retain the draft Hugging Face links and their respective seven- or five-token settings.
  • Correct the DSpark block-drafting explanation and include the Kimi-K3 tokenizer mode and automatic tool-selection option in the serving examples.
  • Add a vllm bench serve command for the configured chat endpoint, with explicit tokenizer, input/output lengths, request count, concurrency, and saved results. Keep the tutorial limited to the standard deployment-template sections.

Does this PR introduce any user-facing change?

No runtime behavior changes. This is a documentation-only update.

How was this patch tested?

  • bash format.sh ci: all hooks passed.
  • bash tools/rtd_build.sh: full English MkDocs build passed in strict mode.
  • PYTHONPATH=. python -m pytest --noconftest -o addopts= -q tests/ut/test_dependency_documentation.py tests/ut/_tools/test_docs_codegen.py tests/ut/_tools/test_generate_zh_docs.py: 25 tests and 19 subtests passed. These documentation-only tests do not require the shared NPU test setup.
  • Checked the rendered tutorial's ten section headings, navigation order, supported-model table columns, and 199 local links/anchors. Parsed all nine shell blocks with bash -n and all five embedded JSON payloads.
  • Verified that all 16 benchmark flags exist in the pinned vLLM command-line parser. The benchmark itself was not run against an NPU service.
  • Verified the three draft ModelScope repositories are accessible, their configs match the Hugging Face versions, and the published weight SHA256 values and sizes match.
  • git diff --check: passed.

No NPU serving, model accuracy, or performance tests were run for this documentation-only change. The hardware topology and capacity settings are retained from the existing full-checkpoint example; the tuning table does not claim a new performance result.

Wire Kimi K3 into the MkDocs navigation and model support matrix. Follow the standard model tutorial structure, document the full-checkpoint deployment and matching installation, and remove development-only checkpoint reduction and regression material.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241
maoxx241 requested review from LCAIZJ and Yikun as code owners August 28, 2026 03:56
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request standardizes the Kimi K3 documentation by aligning it with the project's established model deployment template. It improves user accessibility by integrating the tutorial into the main navigation and supported-model matrix, while providing comprehensive, production-ready guidance for Atlas A3 deployments.

Highlights

  • Tutorial Standardization: Restructured the Kimi K3 tutorial to align with the standard ten-section model deployment template.
  • Documentation Integration: Updated the MkDocs navigation and the supported-model matrix to include Kimi K3.
  • Deployment Guidance: Added detailed Atlas A3 deployment instructions, including serving parameters, DSpark configuration, and multi-node setup.
  • Cleanup: Removed outdated development-only regression checks and checkpoint-reduction instructions.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the Kimi-K3 model documentation to align with the standard deployment template, restructuring the tutorial, adding the model to the supported models matrix, and updating the MkDocs navigation. The reviewer provided a suggested PR title and summary formatted according to the repository's style guide.

Comment thread docs/hooks/nav_titles.py
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 28, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

Use the Eco-Tech W4A8 model link and document both Inferact MLA DSpark checkpoints. Keep the model path and five- or seven-token setting consistent across prerequisites, deployment, tuning, and FAQ.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Comment thread docs/source/tutorials/models/Kimi-K3.md Outdated
Comment thread docs/source/tutorials/models/Kimi-K3.md
Comment thread docs/source/tutorials/models/Kimi-K3.md Outdated
Comment thread docs/source/tutorials/models/Kimi-K3.md
Add verified ModelScope mirrors for all DSpark drafts, provide a serving benchmark example, clarify the Atlas A3 installation scope, and use the consistent Kimi-K3 model name.

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
@maoxx241 maoxx241 changed the title [Doc][Model] Align Kimi K3 tutorial with the deployment template [Doc][Misc] Align Kimi-K3 tutorial with the deployment template Aug 28, 2026
@liqian79
liqian79 merged commit e98b375 into vllm-project:main Aug 28, 2026
17 checks passed
Lethobenthos20 pushed a commit to Lethobenthos20/vllm-ascend that referenced this pull request Sep 4, 2026
…-project#15186)

### What this PR does / why we need it?

Follow up on vllm-project#14454 to align the Kimi-K3 documentation with the model
deployment tutorial template and the existing Kimi, DeepSeek, and Qwen
tutorials.

- Add Kimi-K3 to the MkDocs model navigation and bilingual title
mapping, and restore the generic model tutorial overview.
- Add a conservative A3 W4A8 entry to the supported-model matrix and
link the tutorial to the shared model/feature guides.
- Organize the tutorial into the standard ten sections, including
prerequisites, installation, full-checkpoint deployment, functional
verification, accuracy/performance evaluation, tuning, and FAQ.
- Document the Atlas A3-only installation scope, main-branch A3 image,
and source-build path using the verified upstream vLLM revision; add key
serving-parameter explanations and a reference configuration table.
- Provide ModelScope links for the Eco-Tech W4A8 target, RadixArk GQA
draft, and both Inferact MLA drafts. Retain the draft Hugging Face links
and their respective seven- or five-token settings.
- Correct the DSpark block-drafting explanation and include the Kimi-K3
tokenizer mode and automatic tool-selection option in the serving
examples.
- Add a `vllm bench serve` command for the configured chat endpoint,
with explicit tokenizer, input/output lengths, request count,
concurrency, and saved results. Keep the tutorial limited to the
standard deployment-template sections.

### Does this PR introduce _any_ user-facing change?

No runtime behavior changes. This is a documentation-only update.

### How was this patch tested?

- `bash format.sh ci`: all hooks passed.
- `bash tools/rtd_build.sh`: full English MkDocs build passed in strict
mode.
- `PYTHONPATH=. python -m pytest --noconftest -o addopts= -q
tests/ut/test_dependency_documentation.py
tests/ut/_tools/test_docs_codegen.py
tests/ut/_tools/test_generate_zh_docs.py`: 25 tests and 19 subtests
passed. These documentation-only tests do not require the shared NPU
test setup.
- Checked the rendered tutorial's ten section headings, navigation
order, supported-model table columns, and 199 local links/anchors.
Parsed all nine shell blocks with `bash -n` and all five embedded JSON
payloads.
- Verified that all 16 benchmark flags exist in the pinned vLLM
command-line parser. The benchmark itself was not run against an NPU
service.
- Verified the three draft ModelScope repositories are accessible, their
configs match the Hugging Face versions, and the published weight SHA256
values and sizes match.
- `git diff --check`: passed.

No NPU serving, model accuracy, or performance tests were run for this
documentation-only change. The hardware topology and capacity settings
are retained from the existing full-checkpoint example; the tuning table
does not claim a new performance result.

- vLLM main:
vllm-project/vllm@ba07e4a

---------

Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants