| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-09-28 | 6.9 kB | |
| v0.18.0 source code.tar.gz | 2026-09-28 | 2.9 MB | |
| v0.18.0 source code.zip | 2026-09-28 | 4.1 MB | |
| Totals: 3 Items | 7.0 MB | 0 | |
What's Changed
🚀 Features
- [Feat]: Support output input logprobs by @RunningLeon in https://github.com/InternLM/lmdeploy/pull/4793
- Expand SM90 quantized GEMMs and unify TurboMind linear execution by @lzhangzz in https://github.com/InternLM/lmdeploy/pull/4943
- [ascend] support glm52 by @wanfengcxz in https://github.com/InternLM/lmdeploy/pull/4928
- Support dflash for qwen3.5 by @RunningLeon in https://github.com/InternLM/lmdeploy/pull/4789
- Move shared expert into the MoE FFN layer; shard dense FFN node-locally in MoE models by @lzhangzz in https://github.com/InternLM/lmdeploy/pull/4973
💥 Improvements
- fix: shard MLA attention under hybrid DP and TP by @CUHKSZzxy in https://github.com/InternLM/lmdeploy/pull/4908
- refactor(serving): share chat serving runner by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4876
- feat(pytorch): prototype TurboMind W4A16 backend for AWQ linear by @qescccczmr in https://github.com/InternLM/lmdeploy/pull/4897
- refactor(pytorch): redesign CacheEngine around plans and allocations by @grimoire in https://github.com/InternLM/lmdeploy/pull/4862
- fix: reject multimodal requests for text models by @CUHKSZzxy in https://github.com/InternLM/lmdeploy/pull/4926
- refactor(pytorch): replace backend operator builders with typed build specs by @grimoire in https://github.com/InternLM/lmdeploy/pull/4890
- Migrate TurboMind to C++20 by @lzhangzz in https://github.com/InternLM/lmdeploy/pull/4946
- feat(pytorch): add piecewise CUDA graph prefill by @grimoire in https://github.com/InternLM/lmdeploy/pull/4895
- feat: add XTuner TileLang sparse MLA backend by @CUHKSZzxy in https://github.com/InternLM/lmdeploy/pull/4904
- feat: add request-only cache usage metric by @grimoire in https://github.com/InternLM/lmdeploy/pull/4940
- refactor: stream tool parameters incrementally by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4802
- feat: support checkpoint-engine weight updates by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4913
- feat: add symmetric-memory LM-head all-gather by @qescccczmr in https://github.com/InternLM/lmdeploy/pull/4915
- Refactor PyTorch scheduler ownership and API boundaries by @grimoire in https://github.com/InternLM/lmdeploy/pull/4921
- feat(gpt-oss): combine xgrammar structural_tag grammar with Harmony parser by @windreamer in https://github.com/InternLM/lmdeploy/pull/4907
- Decouple GEMM workspace from Linear handle by @lzhangzz in https://github.com/InternLM/lmdeploy/pull/4978
- feat(kv_connector): mooncake store support mtp by @caikun-pjlab in https://github.com/InternLM/lmdeploy/pull/4948
- fix(pytorch): avoid health RPC scheduling delay by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4966
- refactor: remove legacy OpenAI API client by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4993
🐞 Bug fixes
- fix: fall back to model generation config by @CUHKSZzxy in https://github.com/InternLM/lmdeploy/pull/4929
- fix: initialize MLA KV-B after online FP8 quantization by @CUHKSZzxy in https://github.com/InternLM/lmdeploy/pull/4925
- fix tool_choice='required' by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4909
- fix: exclude glm-5.2 mtp projection from online fp8 by @CUHKSZzxy in https://github.com/InternLM/lmdeploy/pull/4945
- Fix TurboMind vision workspace overflows by @irexyc in https://github.com/InternLM/lmdeploy/pull/4924
- fix(proxy): interpolate node_url in the terminate_node failure response by @Anai-Guo in https://github.com/InternLM/lmdeploy/pull/4941
- fix(disagg): make zmq_disconnect synchronous so p2p_drop_connect actually closes the sockets by @Anai-Guo in https://github.com/InternLM/lmdeploy/pull/4916
- Remove cuBLAS grouped GEMM and give cuBLAS dense priority over native BF16/FP16 kernels by @lzhangzz in https://github.com/InternLM/lmdeploy/pull/4975
- fix(internvit): allocate pinned input buffers on demand by @irexyc in https://github.com/InternLM/lmdeploy/pull/4982
- fix(serve): tolerate messages without a content key in multimodal preprocessing by @matrix72c in https://github.com/InternLM/lmdeploy/pull/4994
- fix: enforce weights-only loading for MemDecode router checkpoints by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4988
- fix(dlinfer): remove CUDA hardcode in NTK rotary embedding for non-CUDA accelerators by @li-lizhe in https://github.com/InternLM/lmdeploy/pull/4986
- fix: build NCCL stub with public device API header by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4998
- fix: correct tool choice handling for Intern-S and GPT-OSS by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4997
- fix: restore getenv after environment parsing errors by @adenzhou1350 in https://github.com/InternLM/lmdeploy/pull/5000
🌐 Other
- [Docs] Fix A100 FP16 benchmark labels by @BingH225 in https://github.com/InternLM/lmdeploy/pull/4919
- fix(deps): pin xgrammar below 0.2.5.post1 to unbreak unit-test CI by @windreamer in https://github.com/InternLM/lmdeploy/pull/4931
- Assert the logprobs count in the restful return checks by @David-Wu1119 in https://github.com/InternLM/lmdeploy/pull/4934
- docs: align docstring parameter names with signatures by @simpleqt in https://github.com/InternLM/lmdeploy/pull/4938
- [Docs] Fix syntax in the Chinese output logits example by @BingH225 in https://github.com/InternLM/lmdeploy/pull/4959
- docs: fix dead heading anchors in CONTRIBUTING and get_started pages by @simpleqt in https://github.com/InternLM/lmdeploy/pull/4937
- improve(autotest): replace APIClient with OpenAI SDK and strict param checks by @littlegy in https://github.com/InternLM/lmdeploy/pull/4887
- ci: add swap for CUDA release builds by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4974
- [ci] refactor(autotest): consolidate layout-based pytest cases by @zhulinJulia24 in https://github.com/InternLM/lmdeploy/pull/4881
- Test: update sleep/wakeup and abort scenarios by @littlegy in https://github.com/InternLM/lmdeploy/pull/4528
- docs: sync ja README quickstart python version with en/zh (3.10 -> 3.12) by @simpleqt in https://github.com/InternLM/lmdeploy/pull/4939
- bump version to v0.18.0 by @lvhan028 in https://github.com/InternLM/lmdeploy/pull/4990
New Contributors
- @BingH225 made their first contribution in https://github.com/InternLM/lmdeploy/pull/4919
- @David-Wu1119 made their first contribution in https://github.com/InternLM/lmdeploy/pull/4934
- @simpleqt made their first contribution in https://github.com/InternLM/lmdeploy/pull/4938
- @li-lizhe made their first contribution in https://github.com/InternLM/lmdeploy/pull/4986
- @adenzhou1350 made their first contribution in https://github.com/InternLM/lmdeploy/pull/5000
Full Changelog: https://github.com/InternLM/lmdeploy/compare/v0.17.0...v0.18.0