Download Latest Version v0.26.0 source code.zip (5.3 MB) Google Add to Preferred Sources
Home / v0.25.2
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-09-24 1.9 kB
v0.25.2 source code.tar.gz 2026-09-24 4.1 MB
v0.25.2 source code.zip 2026-09-24 5.2 MB
Totals: 3 Items   9.3 MB 1

Overview

A small point release focused on backend improvements: a new CUDA convolution kernel with an implicit-GEMM fast path, Vulkan fixes and tuning (misalignment handling in convolution shaders and cooperative-matrix support for Adreno GPUs), an optimized DP4A binary kernel for OpenCL Q6_K GEMM, and a precision guard in the Hexagon backend.

Backend changes

CUDA

Vulkan

OpenCL

Hexagon

More info

Changelog since v0.25.1

a0d15e24 ggml : bump version to 0.25.2 (#1642) c5bf7e20 sync : llama.cpp 10e36b6a vulkan: handle misalignment in conv_2d and conv_3d (llama/29365) 0d1fb478 vulkan: tune KHR cooperative matrix support for Adreno GPUs (llama/29328) 749a9ebb sync : llama.cpp aebfeefb cuda : add conv3d with implicit GEMM (llama/29137) 9a7b3c1e hexagon: reject MUL_MAT_ID when src1 precision is F32 (llama/29348) 600ec61f opencl: add A8 Q6_K non-MoE dp4a binary kernel (llama/29057) 64302f42 scripts : make-release-desc - link previous release in changelog title (#1638)

Source: README.md, updated 2026-09-24