| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-04-23 | 2.9 kB | |
| v0.5.5 source code.tar.gz | 2026-04-23 | 49.7 MB | |
| v0.5.5 source code.zip | 2026-04-23 | 50.1 MB | |
| Totals: 3 Items | 99.8 MB | 0 | |
- 支持 LLaDA2.1-mini、LLaDA2.1-flash 模型,是赤兔首次支持扩散语言模型。
- 支持 Kimi-K2.6。
- 向后兼容 Qwen3-Coder-Next-FP8、GLM4.7-FP8。
- 进一步优化 DeepSeek-V3.2 及类似模型中的稀疏 attention。
- 优化 Qwen3-Next 和 Qwen3.5 模型开启 MTP 时的性能。
- 面向前缀缓存优化 DP 路由策略。
- 优化模型加载速度。
- 工具调用兼容 PD 分离。
- 支持单独控制各模块的日志级别(文档)。
- 将 mooncake 作为 pip 安装时的可选依赖,不再需要单独安装。
- 更新 flashinfer 可选依赖。
- 修复在海光平台上的一些兼容问题。
- 修复
infer.prefix_chunk_size较大时的溢出问题。 - 修复多 stream 导致的显存不能及时释放的问题。
- 修复只有 PCIe 互联的环境上的 allreduce 性能。
- 删除了
infer.use_cuda_graph=auto时在昇腾平台上默认关闭 graph 的一项过时判断。 - 重构算子分发逻辑。
- 重构异步调度。
- 改进代码仓库中的若干测试。
- Added support for LLaDA2.1-mini and LLaDA2.1-flash models, first supported diffusion LLM models.
- Added support for Kimi-K2.6.
- Added backward support for Qwen3-Coder-Next-FP8 and GLM4.7-FP8.
- Further optimized sparse attention in DeepSeek-V3.2 and similar models.
- Optimized Qwen3-Next and Qwen3.5 models when enabling MTP.
- Optimized DP routing strategy with respect to prefix caching.
- Opitmized model loading speed.
- Made tool calling compatible with PD-disaggregation.
- Supported logging level control on specific modules (doc).
- Added mooncake as an optional dependency during pip install. It no longer needed to be installed manually.
- Upgraded flashinfer optional dependency.
- Fixed some compatible issues on Hygon platform.
- Fixed overflow issues when
infer.prefix_chunk_sizeis high. - Fixed late memory freeing caused by multi-streaming.
- Fixed allreduce performance on platforms where PCIe is the only interconnect.
- Removed an out-dated default behaviour that turns off the graph when
infer.use_cuda_graph=auto. - Refactored operator dispatching.
- Refactored asynchornous scheduling.
- Improved multiple tests in the repository.
Official Docker images / 官方 docker 镜像:
- 英伟达 / NVIDIA (arch 8.0, 8.9): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-nvidia_arch_80_89:v0.5.5
- 英伟达 / NVIDIA (arch 9.0): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-nvidia_arch_90:v0.5.5
- 沐曦 / MetaX: qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-muxi:v0.5.5
- 昇腾 / Ascend (A2): qingcheng-ai-cn-beijing.cr.volces.com/public/chitu-ascend_a2:v0.5.5