Download Latest Version v0.26.0 source code.zip (5.3 MB) Google Add to Preferred Sources
Home / v0.25.1
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-09-23 2.4 kB
v0.25.1 source code.tar.gz 2026-09-23 4.1 MB
v0.25.1 source code.zip 2026-09-23 5.2 MB
Totals: 3 Items   9.3 MB 0

Overview

A hotfix release focused on the CUDA backend: sparse flash attention, which had been disabled due to a batch-dependent gate, is re-enabled for long-context prefill (up to ~1.4x prefill speedup at 131k context), with the supporting sparse mask scan kernel significantly sped up. Additionally, Metal gains the missing f32 × bf16 matmul kernels (fixing depthwise convolutions with bf16 weights) and its flash-attention tuning tables are now keyed by GPU family instead of individual SKU, while Vulkan adds MMQ/MMV matmul kernels for the IQ4_XS quantization type.

Backend changes

CUDA

  • Re-enable sparse flash attention for long-context prefill (switched off by llama.cpp#28770 due to a batch-dependent gate), with up to ~1.4x prefill speedup at 131k context (llama.cpp#29298)
  • Speed up the sparse mask scan kernel via compile-time loop unrolling and host-side bound selection (586 → 244 us for the batched sparse op at 49k context) (llama.cpp#29298)

Metal

  • Add the missing f32 × bf16 mul_mv kernel variants, fixing depthwise convolutions with bf16 weights (llama.cpp#28741)
  • Key the flash-attention vector tuned tables by GPU family instead of individual SKU, falling back to the baseline for untuned families and substantially shrinking the tuning tables (llama.cpp#29075)

Vulkan

  • Add MMQ/MMV matmul kernels for the IQ4_XS quantization type (llama.cpp#28415)

More info

Changelog since v0.25.0

e565a8f4 ggml : bump version to 0.25.1 (#1637) ef97dbf9 sync : llama.cpp 2b3b7c73 CUDA: add a reserve to avoid spurious warning on older GCC builds (llama/29317) 60e21b46 metal: add the missing f32 x bf16 mul_mv variants (llama/28741) 4506ab1d CUDA: enable sparse-fa for dsv4 prefill (again) (llama/29298) a70078ad metal : key the fa-vec tuned table by family instead of SKU (llama/29075) 8b93b8f0 vulkan: add IQ4_XS MMQ/MMV matmul kernels (llama/28415)

Source: README.md, updated 2026-09-23