| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-09-24 | 1.9 kB | |
| v0.25.2 source code.tar.gz | 2026-09-24 | 4.1 MB | |
| v0.25.2 source code.zip | 2026-09-24 | 5.2 MB | |
| Totals: 3 Items | 9.3 MB | 1 | |
Overview
A small point release focused on backend improvements: a new CUDA convolution kernel with an implicit-GEMM fast path, Vulkan fixes and tuning (misalignment handling in convolution shaders and cooperative-matrix support for Adreno GPUs), an optimized DP4A binary kernel for OpenCL Q6_K GEMM, and a precision guard in the Hexagon backend.
Backend changes
CUDA
- added
conv3dkernel with an implicit-GEMM fast path (F16 weights) and a direct fallback for F32 shapes (https://github.com/ggml-org/llama.cpp/pull/29137)
Vulkan
- handle misalignment in the
conv_2dandconv_3dmatrix-multiply shaders (https://github.com/ggml-org/llama.cpp/pull/29365) - enable and tune
VK_KHR_cooperative_matrixsupport for Adreno GPUs with hardware matrix cores (https://github.com/ggml-org/llama.cpp/pull/29328)
OpenCL
- added optimized DP4A binary kernel for Q6_K non-MoE GEMM (https://github.com/ggml-org/llama.cpp/pull/29057)
Hexagon
- reject
MUL_MAT_IDwhen thesrc1precision is F32 (not supported) (https://github.com/ggml-org/llama.cpp/pull/29348)
More info
Changelog since v0.25.1
a0d15e24 ggml : bump version to 0.25.2 (#1642) c5bf7e20 sync : llama.cpp 10e36b6a vulkan: handle misalignment in conv_2d and conv_3d (llama/29365) 0d1fb478 vulkan: tune KHR cooperative matrix support for Adreno GPUs (llama/29328) 749a9ebb sync : llama.cpp aebfeefb cuda : add conv3d with implicit GEMM (llama/29137) 9a7b3c1e hexagon: reject MUL_MAT_ID when src1 precision is F32 (llama/29348) 600ec61f opencl: add A8 Q6_K non-MoE dp4a binary kernel (llama/29057) 64302f42 scripts : make-release-desc - link previous release in changelog title (#1638)