Repository navigation
[Doc][Misc] Align Kimi-K3 tutorial with the deployment template - #15186
Conversation
Wire Kimi K3 into the MkDocs navigation and model support matrix. Follow the standard model tutorial structure, document the full-checkpoint deployment and matching installation, and remove development-only checkpoint reduction and regression material. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request standardizes the Kimi K3 documentation by aligning it with the project's established model deployment template. It improves user accessibility by integrating the tutorial into the main navigation and supported-model matrix, while providing comprehensive, production-ready guidance for Atlas A3 deployments. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request updates the Kimi-K3 model documentation to align with the standard deployment template, restructuring the tutorial, adding the model to the supported models matrix, and updating the MkDocs navigation. The reviewer provided a suggested PR title and summary formatted according to the repository's style guide.
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Use the Eco-Tech W4A8 model link and document both Inferact MLA DSpark checkpoints. Keep the model path and five- or seven-token setting consistent across prerequisites, deployment, tuning, and FAQ. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
Add verified ModelScope mirrors for all DSpark drafts, provide a serving benchmark example, clarify the Atlas A3 installation scope, and use the consistent Kimi-K3 model name. Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
…-project#15186) ### What this PR does / why we need it? Follow up on vllm-project#14454 to align the Kimi-K3 documentation with the model deployment tutorial template and the existing Kimi, DeepSeek, and Qwen tutorials. - Add Kimi-K3 to the MkDocs model navigation and bilingual title mapping, and restore the generic model tutorial overview. - Add a conservative A3 W4A8 entry to the supported-model matrix and link the tutorial to the shared model/feature guides. - Organize the tutorial into the standard ten sections, including prerequisites, installation, full-checkpoint deployment, functional verification, accuracy/performance evaluation, tuning, and FAQ. - Document the Atlas A3-only installation scope, main-branch A3 image, and source-build path using the verified upstream vLLM revision; add key serving-parameter explanations and a reference configuration table. - Provide ModelScope links for the Eco-Tech W4A8 target, RadixArk GQA draft, and both Inferact MLA drafts. Retain the draft Hugging Face links and their respective seven- or five-token settings. - Correct the DSpark block-drafting explanation and include the Kimi-K3 tokenizer mode and automatic tool-selection option in the serving examples. - Add a `vllm bench serve` command for the configured chat endpoint, with explicit tokenizer, input/output lengths, request count, concurrency, and saved results. Keep the tutorial limited to the standard deployment-template sections. ### Does this PR introduce _any_ user-facing change? No runtime behavior changes. This is a documentation-only update. ### How was this patch tested? - `bash format.sh ci`: all hooks passed. - `bash tools/rtd_build.sh`: full English MkDocs build passed in strict mode. - `PYTHONPATH=. python -m pytest --noconftest -o addopts= -q tests/ut/test_dependency_documentation.py tests/ut/_tools/test_docs_codegen.py tests/ut/_tools/test_generate_zh_docs.py`: 25 tests and 19 subtests passed. These documentation-only tests do not require the shared NPU test setup. - Checked the rendered tutorial's ten section headings, navigation order, supported-model table columns, and 199 local links/anchors. Parsed all nine shell blocks with `bash -n` and all five embedded JSON payloads. - Verified that all 16 benchmark flags exist in the pinned vLLM command-line parser. The benchmark itself was not run against an NPU service. - Verified the three draft ModelScope repositories are accessible, their configs match the Hugging Face versions, and the published weight SHA256 values and sizes match. - `git diff --check`: passed. No NPU serving, model accuracy, or performance tests were run for this documentation-only change. The hardware topology and capacity settings are retained from the existing full-checkpoint example; the tuning table does not claim a new performance result. - vLLM main: vllm-project/vllm@ba07e4a --------- Signed-off-by: maoxx241 <maomaoyu870@gmail.com>
What this PR does / why we need it?
Follow up on #14454 to align the Kimi-K3 documentation with the model deployment tutorial template and the existing Kimi, DeepSeek, and Qwen tutorials.
vllm bench servecommand for the configured chat endpoint, with explicit tokenizer, input/output lengths, request count, concurrency, and saved results. Keep the tutorial limited to the standard deployment-template sections.Does this PR introduce any user-facing change?
No runtime behavior changes. This is a documentation-only update.
How was this patch tested?
bash format.sh ci: all hooks passed.bash tools/rtd_build.sh: full English MkDocs build passed in strict mode.PYTHONPATH=. python -m pytest --noconftest -o addopts= -q tests/ut/test_dependency_documentation.py tests/ut/_tools/test_docs_codegen.py tests/ut/_tools/test_generate_zh_docs.py: 25 tests and 19 subtests passed. These documentation-only tests do not require the shared NPU test setup.bash -nand all five embedded JSON payloads.git diff --check: passed.No NPU serving, model accuracy, or performance tests were run for this documentation-only change. The hardware topology and capacity settings are retained from the existing full-checkpoint example; the tuning table does not claim a new performance result.