Download Latest Version v0.6.0 source code.zip (50.4 MB) Google Add to Preferred Sources
Home / v0.5.6
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-05-21 4.4 kB
v0.5.6 source code.tar.gz 2026-05-21 49.8 MB
v0.5.6 source code.zip 2026-05-21 50.2 MB
Totals: 3 Items   100.0 MB 0
  • 新模型:Qwen3.6-27B、Qwen3.6-27B-FP8、Qwen3.6-35B-A3B、Qwen3.6-35B-A3B-FP8、DeepSeek-V4-Flash-Base(仅实验性)、DeepSeek-V4-Flash-FP8(仅实验性)。
  • 选择模型时,所有“Qwen3_5”字样变更为“Qwen3.5“。
  • 工具调用请求对可用工具的描述中,“required” 字段不再为必填项。
  • 赤兔中内建的大部分 CUDA 算子支持海光平台。
  • 集成海光平台的多个新算子。
  • infer.fuse_shared_experts 选项兼容 EP。
  • 令 DeepSeek-V3.2 等模型中的 indexer cache 也能接受前缀缓存。
  • GLM-5/5.1-FP8 兼容 infer.mla_absorb=absorb-without-precomp 选项。
  • MTP 兼容 infer.schedule_overlap 优化。
  • (TP×DP)+(ETP×EP) 混合并行。
  • 减少 DeepEP + DeepGEMM Prefill 中的 CPU-GPU 同步。
  • 面向前缀缓存命中率优化 DP rank 和多 instance 间的请求路由。
  • 利用多 kernel 联合 autotune 优化 triton MoE 实现。
  • 通过为不同类型的 KV cache 设置不同 page size 优化 DeepSeek-V3.2 及类似模型。
  • 优化前缀缓存的查询时间。
  • 增加 CHITU_LOG_STACK_TRACE 调试用环境变量。
  • 向请求返回前缀缓存命中统计信息。
  • 内建 profiling 功能支持 PD 分离时 profile P (prefill) instance。
  • 清除实时监控中的部分无效信息。
  • 修复若干 prefill chunk size 较大时的 int32 越界问题。
  • 修复 DeepSeek-V3.2 等模型中 indexer 模块的软 FP8 实现。
  • 修复 PD 分离时 Prefill 结点采用 DP 并行时可能导致的死锁问题。
  • 修复 scheduler 中状态不一致可能导致的死锁问题。
  • 清除 FlashMLA 集成中引入的不必要 padding。
  • 修复 Qwen3.5/3.6 开启 PP 时的一处问题。
  • 修复 warmup 期间 DeepSeek-V3.2 等模型中 indexer KV cache 占用过大的问题。

  • New models: Qwen3.6-27B, Qwen3.6-27B-FP8, Qwen3.6-35B-A3B, Qwen3.6-35B-A3B-FP8, DeepSeek-V4-Flash-Base (experimental only)、DeepSeek-V4-Flash-FP8 (experimental only).
  • All "Qwen3_5" in model names for selection are now changed to "Qwen3.5".
  • The "required" field for parameters in tool descriptions in tool-calling requests are no longer required.
  • Most of the built-in CUDA kernels in Chitu now supports Hygon platforms.
  • New integration for multiple new operators on Hygon platforms.
  • infer.fuse_shared_experts argument is now compatible with EP.
  • Indexer cache in DeepSeek-V3.2 and similar models can now be cached by prefix cache.
  • GLM-5/5.1-FP8 is now compatible with infer.mla_absorb=absorb-without-precomp argument.
  • MTP is now compatible with infer.schedule_overlap optimization.
  • (TP×DP)+(ETP×EP) hybrid parallelism.
  • Fewer CPU-GPU synchronizations in DeepEP + DeepGEMM prefill steps.
  • Optimization on prefix cache hit rate on request routing among DP ranks and instances.
  • Optimization on triton MoE implementations via co-auto-tune on multiple kernels.
  • Optimization on DeepSeek-V3.2 and similar models via setting different page sizes for different types of KV caches.
  • Optimization on prefix cache query time.
  • New CHITU_LOG_STACK_TRACE environment variable for debugging.
  • Add prefix cache hit status to responses.
  • The built-in profiling feature now support profiling the P (prefill) instances in PD-disaggregated deployment.
  • Less redundant data in monitored metrics.
  • Fixed several issues on int32 overflow on large prefill chunk sizes.
  • Fixed the soft FP8 implementation on indexer module in DeepSeek-V3.2 and similar models.
  • Fixed a dead-lock issue on PD-disaggregation when prefill instances use DP.
  • Fixed a dead-lock issue caused by inconsistent states in scheduler.
  • Removed some redundant paddings from MLA integration.
  • Fixed an issue on Qwen3.5/3.6 models when using PP.
  • Fixed unexpected large indexer KV cache size during warming-up in DeepSeek-V3.2 and similar models.

Official Docker images / 官方 docker 镜像:

Source: README.md, updated 2026-05-21