Download Latest Version v0.26.0 source code.zip (5.3 MB) Google Add to Preferred Sources
Home / v0.25.0
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-09-23 17.5 kB
v0.25.0 source code.tar.gz 2026-09-23 4.1 MB
v0.25.0 source code.zip 2026-09-23 5.2 MB
Totals: 3 Items   9.3 MB 0

Overview

This release expands hyper-connection, flash-attention, and fused MoE/SSM support across CPU, GPU, and accelerator backends. It also improves backend robustness, quantization, data-layout handling, and RPC/meta buffer management. Numerous correctness fixes and performance tuning land across all major backends, including new kernels, fusions, and op coverage.

API changes

Core changes

  • Added gated hyper-connection preprocessing (llama/28901).
  • Added hc_post identity path without comb (llama/28901).
  • Fixed dimension/stride truncation in ggml_permute (llama/29227).
  • Scheduler propagates graph reservation failures (llama/26070).
  • Optimized IQ1_M with per-block prefix sums (llama/28706).
  • GGUF data alignment is now relative to GGUF start (llama/28993).

Backend changes

CPU

CUDA / HIP

Metal

  • Added MoE and SSM_CONV fusion optimizations (llama/28948).
  • Added hyper-connection fusion optimizations (llama/28948).
  • Added arbitrary hc support in dsv4_hc_pre (llama/29169).
  • Added new hyper-connection ops (llama/29000).
  • Added F16 input support to FWHT (llama/29094).
  • Added flash attention kernels for new head sizes (llama/28599).
  • Fixed flash attention mask bounds (llama/29220).
  • Fixed flash attention support checks (llama/29122).
  • Fixed mul_mm_id NaN for large F16 activations (llama/26223).
  • Gated mul_mm_id src1 rescale behind ggml_prec (llama/29029).
  • Fixed macOS SDK deprecation warnings (llama/29136).
  • Simplified fusion-pattern declarations (llama/29206).

Vulkan

  • Added IQ3_S MMQ matmul kernels (llama/28822).
  • Added sparse flash attention support (llama/28105).
  • Added Intel Xe flash attention optimization kernels (llama/24406).
  • Added new hyper-connection ops (llama/28988).
  • Refactored Vulkan buffers/debug into shared modules (llama/28732).
  • Raised mul_mat_id row-id limit to 512 experts (llama/28501).
  • Skipped unneeded MoE work in the coopmat1 path (llama/25483).
  • Fixed buffer reference alignment in im2col shaders (llama/28996).
  • Worked around the NV argsort_large compiler bug (llama/28975).
  • Made MUL_MAT_ID BN/2 tail unconditional (llama/28923).

OpenCL

SYCL

  • Added GET_ROWS_BACK support for FP32/FP16 (llama/25266).
  • Added gated dsv4_hc_pre support (llama/29132).
  • Added optional hc_post comb matrix support (llama/29132).
  • Added MMVQ GLU fusion (llama/28931).
  • Added rms_norm+scale and ssm_conv+silu fusions (llama/28931).
  • Fused the SiLU epilogue into the ssm_conv kernel (llama/28929).
  • Improved MKL-FA softmax load coalescing (llama/28918).
  • Fixed pinned memory to use the right device context (llama/28895).
  • Fixed a large-allocation memory error (llama/28953).
  • Fixed the split buffer type function signature (llama/28981).
  • Added hadamard fp16 unit-test support (llama/29218).

Hexagon

WebGPU

OpenVINO

  • Updated OpenVINO to 2026.4 (llama/29009).
  • Fixed clangd and MSVC warnings (llama/29009).
  • Optimized stateful decode and GPU MoE inference (llama/28638).
  • Added fused compressed-MoE and KV-state passes (llama/28638).
  • Expanded op coverage for rope, norms, MoE, and flash-attention graphs.

RPC / Meta

More info

Changelog since v0.24.0

cae37675 ggml : bump version to 0.25.0 (#1635) d49e952b common : fix for two functions when top_k exceeds the vocabulary size. (#1633) 07a4b1ce sycl : fix compile warnings 819fe452 sync : llama.cpp 51404d30 ggml-meta: resolve multi buffer views (llama/29266) d5afa81e cuda: top-k MoE should always fire (llama/28432) 21b34ca9 sycl : support new UT case for mul_mat_hadamard fp16 (llama/29218) dd666be0 sycl: extend MMVQ GLU fusion, add rms_norm+scale and ssm_conv+silu fusions (llama/28931) b84485e2 sycl : support op get_rows_back, only support fp32/fp16 (llama/25266) 42410209 vulkan: hide internal symbols to prevent duplicate-dlopen state destruction (llama/29139) 4954fc33 hex-dma: introduce direct-mapped DMA cache that is better suited for HVX FA mask handling (llama/29282) 45d3afd1 HIP : optimize IQ2/IQ3 (__vsub4 __vcmpne4) using SWAR (llama/27962) f92e0376 opencl: add bin kernel kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin (llama/29056) 036edabe vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LPG Plus/Xe2/Xe3) (llama/24406) 9a67799f metal : gate mul_mm_id src1 rescale behind ggml_prec (llama/29029) 2def236a ggml : IQ1_M build prefix sums once per block (llama/28706) cd9cc61f Performance tune for gemma4-26b-a4b flash attention shape. (llama/28450) 179b60f2 sync : llama.cpp 7f6fffca opencl: add A8 Q4_0 non-MoE dp4a binary kernel (llama/29055) 6574befe hexagon: new HMX-optimized GATED_DELTA_NET (llama/29199) fee3f947 sync : llama.cpp 5da72455 metal : fix mask bounds in flash attention block pre-pass (llama/29220) 6f74a142 cuda: fix sm_70 tile compilation error (llama/29224) 3847332e ggml-cuda : convert contiguous tensors four elements at a time (llama/29155) a6e77bf4 cuda : accelerate conv2d with implicit GEMM (llama/29135) 892cd446 ggml : fix dimension and stride truncation in ggml_permute (llama/29227) 4de56aa0 sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix (llama/29132) 95756e08 sycl : pinned memory use right device context instead of 0 (llama/28895) 4a77ed32 CUDA: Follow up of [#25635], refactoring FA shared smem swizzle (llama/28536) cf72a89e tests/test-backend-ops : allow regex entries in the -o filter (llama/29204) 9fab492c tests : remove stale comment (llama/29140) 88d2340a ggml-metal : simplify fusion pattern op list declaration (llama/29206) 0187e859 sycl : coalesce MKL-FA softmax loads instead of one work-item per row (llama/28918) e1f0205a ggml-cpu: ARM Repack kernels for Q1_0 (llama/23492) 18194b98 hexagon: overhaul of buffer and DMA handling to support 64bit mappings + improvements (llama/29197) 67ecd06d cuda : tune MMVQ to MMQ crossover for SM70 (Volta) (llama/28912) af5704b7 metal : fix deprecation warnings from macOS 27 SDK (llama/29136) 5d56a164 webgpu : add fused gdn + cpy (llama/28976) 483962ab CUDA: tune FA for Gemma 4 on Ampere or newer (llama/29152) d6362c19 metal : support arbitrary hc in dsv4_hc_pre (llama/29169) ec545717 CUDA: enable sparse fa for qwen4 (llama/28770) 40c4b204 metal: add F16 input to the FWHT (llama/29094) 695046b7 hexagon: enable I32 GET_ROWS (llama/29116) 12434c59 hexagon: add support for GEGLU_QUICK (llama/29114) 196c98ab hexagon: enable support for TOP_K op (llama/29113) d71b6d7b metal : add MoE and SSM_CONV fusion optimizations (llama/28948) 5890241d metal : fix FA support checks (llama/29122) ff93f711 metal : support qwen4exp hc ops (llama/29000) da963d58 cuda : fix CUB argsort corruption caused by in-place keys (llama/28389) fa2f6ba0 opencl: add support for bin kernel flash_attn_f32_f16_bin (llama/29046) f5513ba8 hexagon: add ROLL op support (llama/29105) b3afd1ea hexagon: im2col update (llama/29103) 24701ef0 hexagon: HMX flash-attention head_dim padding (support DK=DV=72) (llama/26539) 175b35ca opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin (llama/28678) 4f37e0e0 ggml-cpu: add F16 input to the FWHT (llama/27779) bd3420dc ggml-webgpu: fix supports_op condition for GET_ROWS (llama/28978) 54610de9 ggml : handle graph buffer reservation failure (llama/26070) 3323448c vulkan: add IQ3_S MMQ matmul kernels (llama/28822) 5dbf6187 vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (llama/28501) 8433dd0f openvino : Update OpenVINO to 2026.4;fix clangd,MSVC warnings; (llama/29009) 3e9c6ec1 vulkan: split buffers and debug code into separate files, add shared headers (llama/28732) 1950eb0e gguf : align the data section relative to the GGUF start, not the file (llama/28993) b2880dbe sycl : fix the B70 mem allocate error when >19.3GB (llama/28953) 2a796566 vulkan: skip unneeded MoE work in mul_mm coopmat1 path (llama/25483) c28a67ca sycl: ssm_conv: fuse the SiLU epilogue into the ssm_conv kernel (llama/28929) 31f3f01a opencl: fix various warnings (llama/28984) 874c7ad7 vulkan: fix buffer_reference alignment in im2col shaders (llama/28996) 49bf19b5 vulkan: support qwen4exp hc ops (llama/28988) bf0c7543 Fix function signature for ggml_backend_sycl_split_buffer_type (llama/28981) 959dc03f vulkan: work around NV bug with argsort_large.comp (llama/28975) 348c9c83 hexagon: Support for K-Quants Q4_K and Q6_K (llama/28994) a9e1a3aa hexagon: accept the zeroed rope probe in supports_op (llama/28995) a8998167 CUDA/HIP: improve access patterns in im2col (llama/28013) 228ac30a spacemit : fix wrong transpose function for int16 data (llama/25161) e6a3d0db rpc : invalidate cached compute graph when a referenced buffer is freed (llama/24292) 7f32c83a qwen4exp: add hc ops (llama/28901) cfd1da6d HIP: broaden MoE ncols_opt tile heuristic on RDNA3.5 architecture (llama/28935) 31119c42 vulkan: make MUL_MAT_ID BN/2 tail unconditional (llama/28923) 9a4a9ea5 metal: fix NaN in mul_mm_id when activations exceed f16 range (llama/26223) e48ccd42 ci : add self-hosted webgpu to hf-jobs (llama/28712) cf54e96f hexagon: add back missing contiguous fast-path and hvx_copy_uu for each run (llama/28886) b325bb2a hex-cpy: use dma if src and dst are contiguous (llama/28906) 9a54ff12 HIP: Enable AllReduce for ROCm (llama/27825) 5fb997bb opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (llama/27637) fd2510e8 rpc : hash-cache only weights (llama/28789) da48a8c8 cuda: support row-contiguous SUM_ROWS (llama/26308) fbbbc117 vulkan: support sparse Flash Attention (llama/28105) 6003c35a OpenVINO: optimize stateful decode and GPU MoE inference (llama/28638) fecab48f opencl: add generic ssm_scan (llama/28881) 05f37575 metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (llama/28599) d390a09b cuda : enable i16 and i32 for DUP (llama/28897) c535c343 HIP: fattn-mma: use fp32 accumulation on MFMA devices (llama/28576)

Source: README.md, updated 2026-09-23