Download Latest Version v0.6.0 source code.zip (50.4 MB) Google Add to Preferred Sources
Home / v0.5.4
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-04-09 3.3 kB
v0.5.4 source code.tar.gz 2026-04-09 49.6 MB
v0.5.4 source code.zip 2026-04-09 50.0 MB
Totals: 3 Items   99.7 MB 0
  • 新模型:GLM-5.1 系列和 Kimi-K2.5。
  • 兼容 OpenAI /response API(文档)。
  • 支持在 Qwen3-Next 和 Qwen3.5 系列模型中使用 MTP。
  • 支持在 DP 并行时例外地 TP 并行 embed_token 和 lm_head 层(设置 infer.embed_tokens_lm_head_tp_size 使用)。
  • 新增优雅关闭服务的功能。
  • 新增内置的 profiling 功能。
  • 优化 DeepSeek-V3.2 及类似模型中的稀疏 attention。
  • 弃用 infer.max_reqs 选项,改为意义更明确的 infer.max_batch_size 和 infer.max_concurrent_requests 选项。
  • 增加 infer.mla_absorb=auto 选项,并将 auto 作为默认值。
  • 重构有状态的采样器。
  • 重构 KV cache,使其更能适配不同种类的模型。
  • 重构代码仓库中包含的若干测试。
  • 修复前缀缓存中的若干问题。
  • 修复 MoE 算子分发以及 DeepGEMM 集成的若干问题。
  • 修复 log 不能被完整存储到文件的问题。
  • 修复带有稠密层的 MoE 模型无法使用 DP+EP+PP 并行的问题。
  • 修复依据 infer.memory_utilization 选项自动分配 KV cache 时的计算错误。
  • 修复监控数据中显存占用量的显示错误。
  • 进一步修复 tokenizer 解码特殊字符时的问题。

  • New models: GLM-5.1 series and Kimi-K2.5.
  • Compatibility with OpenAI /response API (doc).
  • MTP support in Qwen3-Next and Qwen3.5 series models.
  • Support of exceptionally using TP for embed_token and lm_head layer when using DP for other layers (set infer.embed_tokens_lm_head_tp_size to use).
  • New feature to gracefully shut down the service.
  • New built-in profiling support.
  • Optimize sparse attention in DeepSeek-V3.2 and similar models.
  • Deprecating infer.max_reqs argument and replacing with more explicit infer.max_batch_size and infer.max_concurrent_requests arguments.
  • New infer.mla_absorb=auto argument, where auto is the new defult value.
  • Refactoration of samplers with states.
  • Refactoration of KV cache to better support different types of models.
  • Refactoration of some tests in the repo.
  • Fixes on prefix caching.
  • Fixes on MoE operator dispatching and DeepGEMM integration.
  • Fix on the issue that logs cannot be fully saved to files.
  • Fix on compatibility between MoE modes with dense layers, and DP+EP+PP parallelism.
  • Fix on the calculation when automatically allocating KV cache with respect to infer.memory_utilization.
  • Fix on device memory usage in statistics.
  • Further fix on tokenizer decoding of special characters.

Official Docker images / 官方 docker 镜像:

Source: README.md, updated 2026-04-09