| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-09-23 | 17.5 kB | |
| v0.25.0 source code.tar.gz | 2026-09-23 | 4.1 MB | |
| v0.25.0 source code.zip | 2026-09-23 | 5.2 MB | |
| Totals: 3 Items | 9.3 MB | 0 | |
Overview
This release expands hyper-connection, flash-attention, and fused MoE/SSM support across CPU, GPU, and accelerator backends. It also improves backend robustness, quantization, data-layout handling, and RPC/meta buffer management. Numerous correctness fixes and performance tuning land across all major backends, including new kernels, fusions, and op coverage.
API changes
- Added gated
ggml_dsv4_hc_pre_gated()(llama/28901). - Allowed
ggml_dsv4_hc_post()NULLcomb (llama/28901). - Bumped the RPC protocol major version to 7 (llama/24292).
Core changes
- Added gated hyper-connection preprocessing (llama/28901).
- Added
hc_postidentity path without comb (llama/28901). - Fixed dimension/stride truncation in
ggml_permute(llama/29227). - Scheduler propagates graph reservation failures (llama/26070).
- Optimized IQ1_M with per-block prefix sums (llama/28706).
- GGUF data alignment is now relative to GGUF start (llama/28993).
Backend changes
CPU
- Added ARM repack kernels for Q1_0 (llama/23492).
- Fixed SpacemiT int16 transpose (llama/25161).
- Added F16 input for FWHT and hadamard
MUL_MAT(llama/27779).
CUDA / HIP
- Added implicit-GEMM
conv2dsupport (llama/29135). - Convert contiguous tensors four elements at a time (llama/29155).
- Added i16/i32
DUPsupport (llama/28897). - Added row-contiguous
SUM_ROWSsupport (llama/26308). - Refactored flash attention shared-memory swizzle (llama/28536).
- Tuned flash attention shapes for Ampere and newer (llama/29152).
- Enabled sparse flash attention (llama/28770).
- Tuned the SM70 MMVQ/MMQ crossover (llama/28912).
- Fixed the sm_70 tile compilation error (llama/29224).
- Fixed CUB argsort corruption from in-place keys (llama/28389).
- Fixed top-k MoE dispatch (llama/28432).
- Improved im2col access patterns (llama/28013).
- HIP: optimized IQ2/IQ3 with SWAR (llama/27962).
- HIP: enabled AllReduce for ROCm (llama/27825).
- HIP: broadened RDNA3.5 MoE tile heuristics (llama/28935).
- HIP: use F32 accumulation for MFMA flash attention (llama/28576).
Metal
- Added MoE and SSM_CONV fusion optimizations (llama/28948).
- Added hyper-connection fusion optimizations (llama/28948).
- Added arbitrary
hcsupport indsv4_hc_pre(llama/29169). - Added new hyper-connection ops (llama/29000).
- Added F16 input support to FWHT (llama/29094).
- Added flash attention kernels for new head sizes (llama/28599).
- Fixed flash attention mask bounds (llama/29220).
- Fixed flash attention support checks (llama/29122).
- Fixed
mul_mm_idNaN for large F16 activations (llama/26223). - Gated
mul_mm_idsrc1 rescale behindggml_prec(llama/29029). - Fixed macOS SDK deprecation warnings (llama/29136).
- Simplified fusion-pattern declarations (llama/29206).
Vulkan
- Added IQ3_S MMQ matmul kernels (llama/28822).
- Added sparse flash attention support (llama/28105).
- Added Intel Xe flash attention optimization kernels (llama/24406).
- Added new hyper-connection ops (llama/28988).
- Refactored Vulkan buffers/debug into shared modules (llama/28732).
- Raised
mul_mat_idrow-id limit to 512 experts (llama/28501). - Skipped unneeded MoE work in the coopmat1 path (llama/25483).
- Fixed buffer reference alignment in im2col shaders (llama/28996).
- Worked around the NV argsort_large compiler bug (llama/28975).
- Made
MUL_MAT_IDBN/2 tail unconditional (llama/28923).
OpenCL
- Added an A8 Q4_0 non-MoE dp4a binary kernel (llama/29055).
- Added an A8 q4_K/q8_1 binary GEMM kernel (llama/29056).
- Added an A8 q6_K binary GEMM kernel (llama/28678).
- Added a binary flash attention kernel (llama/29046).
- Added generic
ssm_scansupport (llama/28881). - Choose the MoE expert matmul by batch size (llama/27637).
SYCL
- Added
GET_ROWS_BACKsupport for FP32/FP16 (llama/25266). - Added gated
dsv4_hc_presupport (llama/29132). - Added optional
hc_postcomb matrix support (llama/29132). - Added MMVQ GLU fusion (llama/28931).
- Added rms_norm+scale and ssm_conv+silu fusions (llama/28931).
- Fused the SiLU epilogue into the ssm_conv kernel (llama/28929).
- Improved MKL-FA softmax load coalescing (llama/28918).
- Fixed pinned memory to use the right device context (llama/28895).
- Fixed a large-allocation memory error (llama/28953).
- Fixed the split buffer type function signature (llama/28981).
- Added hadamard fp16 unit-test support (llama/29218).
Hexagon
- Added an HMX-optimized gated delta net (llama/29199).
- Added Q4_K and Q6_K K-quants support (llama/28994).
- Added
ROLLop support (llama/29105). - Added
TOP_Kop support (llama/29113). - Added
GEGLU_QUICKop support (llama/29114). - Added I32
GET_ROWSsupport (llama/29116). - Overhauled buffer/DMA handling for 64-bit mappings (llama/29197).
- Added direct-mapped DMA cache for FA masks (llama/29282).
- Updated im2col support (llama/29103).
- Added flash attention head-dim padding (llama/26539).
- Restored contiguous fast-paths and
hvx_copy_uu(llama/28886). - Use DMA for contiguous hex-cpy (llama/28906).
- Accept the zeroed rope probe in
supports_op(llama/28995).
WebGPU
- Added fused gated-delta-net + copy (llama/28976).
- Fixed
GET_ROWSsupport checks (llama/28978). - Added self-hosted WebGPU CI coverage (llama/28712).
OpenVINO
- Updated OpenVINO to 2026.4 (llama/29009).
- Fixed clangd and MSVC warnings (llama/29009).
- Optimized stateful decode and GPU MoE inference (llama/28638).
- Added fused compressed-MoE and KV-state passes (llama/28638).
- Expanded op coverage for rope, norms, MoE, and flash-attention graphs.
RPC / Meta
- RPC invalidates cached graphs on buffer free (llama/24292).
- RPC hashes only weights (llama/28789).
- Meta backend resolves multi-buffer views (llama/29266).
More info
Changelog since v0.24.0
cae37675 ggml : bump version to 0.25.0 (#1635)
d49e952b common : fix for two functions when top_k exceeds the vocabulary size. (#1633)
07a4b1ce sycl : fix compile warnings
819fe452 sync : llama.cpp
51404d30 ggml-meta: resolve multi buffer views (llama/29266)
d5afa81e cuda: top-k MoE should always fire (llama/28432)
21b34ca9 sycl : support new UT case for mul_mat_hadamard fp16 (llama/29218)
dd666be0 sycl: extend MMVQ GLU fusion, add rms_norm+scale and ssm_conv+silu fusions (llama/28931)
b84485e2 sycl : support op get_rows_back, only support fp32/fp16 (llama/25266)
42410209 vulkan: hide internal symbols to prevent duplicate-dlopen state destruction (llama/29139)
4954fc33 hex-dma: introduce direct-mapped DMA cache that is better suited for HVX FA mask handling (llama/29282)
45d3afd1 HIP : optimize IQ2/IQ3 (__vsub4 __vcmpne4) using SWAR (llama/27962)
f92e0376 opencl: add bin kernel kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin (llama/29056)
036edabe vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LPG Plus/Xe2/Xe3) (llama/24406)
9a67799f metal : gate mul_mm_id src1 rescale behind ggml_prec (llama/29029)
2def236a ggml : IQ1_M build prefix sums once per block (llama/28706)
cd9cc61f Performance tune for gemma4-26b-a4b flash attention shape. (llama/28450)
179b60f2 sync : llama.cpp
7f6fffca opencl: add A8 Q4_0 non-MoE dp4a binary kernel (llama/29055)
6574befe hexagon: new HMX-optimized GATED_DELTA_NET (llama/29199)
fee3f947 sync : llama.cpp
5da72455 metal : fix mask bounds in flash attention block pre-pass (llama/29220)
6f74a142 cuda: fix sm_70 tile compilation error (llama/29224)
3847332e ggml-cuda : convert contiguous tensors four elements at a time (llama/29155)
a6e77bf4 cuda : accelerate conv2d with implicit GEMM (llama/29135)
892cd446 ggml : fix dimension and stride truncation in ggml_permute (llama/29227)
4de56aa0 sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix (llama/29132)
95756e08 sycl : pinned memory use right device context instead of 0 (llama/28895)
4a77ed32 CUDA: Follow up of [#25635], refactoring FA shared smem swizzle (llama/28536)
cf72a89e tests/test-backend-ops : allow regex entries in the -o filter (llama/29204)
9fab492c tests : remove stale comment (llama/29140)
88d2340a ggml-metal : simplify fusion pattern op list declaration (llama/29206)
0187e859 sycl : coalesce MKL-FA softmax loads instead of one work-item per row (llama/28918)
e1f0205a ggml-cpu: ARM Repack kernels for Q1_0 (llama/23492)
18194b98 hexagon: overhaul of buffer and DMA handling to support 64bit mappings + improvements (llama/29197)
67ecd06d cuda : tune MMVQ to MMQ crossover for SM70 (Volta) (llama/28912)
af5704b7 metal : fix deprecation warnings from macOS 27 SDK (llama/29136)
5d56a164 webgpu : add fused gdn + cpy (llama/28976)
483962ab CUDA: tune FA for Gemma 4 on Ampere or newer (llama/29152)
d6362c19 metal : support arbitrary hc in dsv4_hc_pre (llama/29169)
ec545717 CUDA: enable sparse fa for qwen4 (llama/28770)
40c4b204 metal: add F16 input to the FWHT (llama/29094)
695046b7 hexagon: enable I32 GET_ROWS (llama/29116)
12434c59 hexagon: add support for GEGLU_QUICK (llama/29114)
196c98ab hexagon: enable support for TOP_K op (llama/29113)
d71b6d7b metal : add MoE and SSM_CONV fusion optimizations (llama/28948)
5890241d metal : fix FA support checks (llama/29122)
ff93f711 metal : support qwen4exp hc ops (llama/29000)
da963d58 cuda : fix CUB argsort corruption caused by in-place keys (llama/28389)
fa2f6ba0 opencl: add support for bin kernel flash_attn_f32_f16_bin (llama/29046)
f5513ba8 hexagon: add ROLL op support (llama/29105)
b3afd1ea hexagon: im2col update (llama/29103)
24701ef0 hexagon: HMX flash-attention head_dim padding (support DK=DV=72) (llama/26539)
175b35ca opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin (llama/28678)
4f37e0e0 ggml-cpu: add F16 input to the FWHT (llama/27779)
bd3420dc ggml-webgpu: fix supports_op condition for GET_ROWS (llama/28978)
54610de9 ggml : handle graph buffer reservation failure (llama/26070)
3323448c vulkan: add IQ3_S MMQ matmul kernels (llama/28822)
5dbf6187 vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (llama/28501)
8433dd0f openvino : Update OpenVINO to 2026.4;fix clangd,MSVC warnings; (llama/29009)
3e9c6ec1 vulkan: split buffers and debug code into separate files, add shared headers (llama/28732)
1950eb0e gguf : align the data section relative to the GGUF start, not the file (llama/28993)
b2880dbe sycl : fix the B70 mem allocate error when >19.3GB (llama/28953)
2a796566 vulkan: skip unneeded MoE work in mul_mm coopmat1 path (llama/25483)
c28a67ca sycl: ssm_conv: fuse the SiLU epilogue into the ssm_conv kernel (llama/28929)
31f3f01a opencl: fix various warnings (llama/28984)
874c7ad7 vulkan: fix buffer_reference alignment in im2col shaders (llama/28996)
49bf19b5 vulkan: support qwen4exp hc ops (llama/28988)
bf0c7543 Fix function signature for ggml_backend_sycl_split_buffer_type (llama/28981)
959dc03f vulkan: work around NV bug with argsort_large.comp (llama/28975)
348c9c83 hexagon: Support for K-Quants Q4_K and Q6_K (llama/28994)
a9e1a3aa hexagon: accept the zeroed rope probe in supports_op (llama/28995)
a8998167 CUDA/HIP: improve access patterns in im2col (llama/28013)
228ac30a spacemit : fix wrong transpose function for int16 data (llama/25161)
e6a3d0db rpc : invalidate cached compute graph when a referenced buffer is freed (llama/24292)
7f32c83a qwen4exp: add hc ops (llama/28901)
cfd1da6d HIP: broaden MoE ncols_opt tile heuristic on RDNA3.5 architecture (llama/28935)
31119c42 vulkan: make MUL_MAT_ID BN/2 tail unconditional (llama/28923)
9a4a9ea5 metal: fix NaN in mul_mm_id when activations exceed f16 range (llama/26223)
e48ccd42 ci : add self-hosted webgpu to hf-jobs (llama/28712)
cf54e96f hexagon: add back missing contiguous fast-path and hvx_copy_uu for each run (llama/28886)
b325bb2a hex-cpy: use dma if src and dst are contiguous (llama/28906)
9a54ff12 HIP: Enable AllReduce for ROCm (llama/27825)
5fb997bb opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (llama/27637)
fd2510e8 rpc : hash-cache only weights (llama/28789)
da48a8c8 cuda: support row-contiguous SUM_ROWS (llama/26308)
fbbbc117 vulkan: support sparse Flash Attention (llama/28105)
6003c35a OpenVINO: optimize stateful decode and GPU MoE inference (llama/28638)
fecab48f opencl: add generic ssm_scan (llama/28881)
05f37575 metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (llama/28599)
d390a09b cuda : enable i16 and i32 for DUP (llama/28897)
c535c343 HIP: fattn-mma: use fp32 accumulation on MFMA devices (llama/28576)