| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-04-09 | 3.3 kB | |
| v0.5.4 source code.tar.gz | 2026-04-09 | 49.6 MB | |
| v0.5.4 source code.zip | 2026-04-09 | 50.0 MB | |
| Totals: 3 Items | 99.7 MB | 0 | |
- 新模型:GLM-5.1 系列和 Kimi-K2.5。
- 兼容 OpenAI
/responseAPI(文档)。 - 支持在 Qwen3-Next 和 Qwen3.5 系列模型中使用 MTP。
- 支持在 DP 并行时例外地 TP 并行
embed_token和lm_head层(设置infer.embed_tokens_lm_head_tp_size使用)。 - 新增优雅关闭服务的功能。
- 新增内置的 profiling 功能。
- 优化 DeepSeek-V3.2 及类似模型中的稀疏 attention。
- 弃用
infer.max_reqs选项,改为意义更明确的infer.max_batch_size和infer.max_concurrent_requests选项。 - 增加
infer.mla_absorb=auto选项,并将auto作为默认值。 - 重构有状态的采样器。
- 重构 KV cache,使其更能适配不同种类的模型。
- 重构代码仓库中包含的若干测试。
- 修复前缀缓存中的若干问题。
- 修复 MoE 算子分发以及 DeepGEMM 集成的若干问题。
- 修复 log 不能被完整存储到文件的问题。
- 修复带有稠密层的 MoE 模型无法使用 DP+EP+PP 并行的问题。
- 修复依据
infer.memory_utilization选项自动分配 KV cache 时的计算错误。 - 修复监控数据中显存占用量的显示错误。
- 进一步修复 tokenizer 解码特殊字符时的问题。
- New models: GLM-5.1 series and Kimi-K2.5.
- Compatibility with OpenAI
/responseAPI (doc). - MTP support in Qwen3-Next and Qwen3.5 series models.
- Support of exceptionally using TP for
embed_tokenandlm_headlayer when using DP for other layers (setinfer.embed_tokens_lm_head_tp_sizeto use). - New feature to gracefully shut down the service.
- New built-in profiling support.
- Optimize sparse attention in DeepSeek-V3.2 and similar models.
- Deprecating
infer.max_reqsargument and replacing with more explicitinfer.max_batch_sizeandinfer.max_concurrent_requestsarguments. - New
infer.mla_absorb=autoargument, whereautois the new defult value. - Refactoration of samplers with states.
- Refactoration of KV cache to better support different types of models.
- Refactoration of some tests in the repo.
- Fixes on prefix caching.
- Fixes on MoE operator dispatching and DeepGEMM integration.
- Fix on the issue that logs cannot be fully saved to files.
- Fix on compatibility between MoE modes with dense layers, and DP+EP+PP parallelism.
- Fix on the calculation when automatically allocating KV cache with respect to
infer.memory_utilization. - Fix on device memory usage in statistics.
- Further fix on tokenizer decoding of special characters.
Official Docker images / 官方 docker 镜像:
- 英伟达 / NVIDIA (arch 8.0, 8.9): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-nvidia_arch_80_89:v0.5.4
- 英伟达 / NVIDIA (arch 9.0): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-nvidia_arch_90:v0.5.4
- 沐曦 / MetaX: qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-muxi:v0.5.4
- 昇腾 / Ascend (A2): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-ascend_a2:v0.5.4
- 昇腾 / Ascend (A3): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-ascend_a3:v0.5.4