Download Latest Version v0.24.0 source code.zip (5.1 MB) Google Add to Preferred Sources
Home / v0.21.0
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-08-21 3.5 kB
v0.21.0 source code.tar.gz 2026-08-21 3.7 MB
v0.21.0 source code.zip 2026-08-21 4.8 MB
Totals: 3 Items   8.5 MB 3

Overview

New version has been released.

More info: dist : releases and versioning of ggml-org projects

Changelog since v0.20.2

8599e0ea sync : llama.cpp 9a538ea5 Revert "sycl : add Q2_K reordered MMVQ and ESIMD kernels (llama/26336)" (llama/27486) 28d7ab84 ci : extract build job into separate release workflow (#1598) 91147d0c ggml : bump version to 0.21.0 (#1597) 19a38580 scripts : restore release.sh preparation script (#1596) bb83ce08 ci : replace release workflow with make-release (#1595) 5da4cf17 ci : split self-hosted jobs into separate workflow (#1594) 33be41c5 sync : llama.cpp 451e766e kleidiai : add SME2 F32 GEMV kernel support (llama/26891) 59bbab9c sycl : add Q2_K reordered MMVQ and ESIMD kernels (llama/26336) 5a232ab4 test : make the FA V-is-view-of-K case a test case parameter (llama/27394) 14dfda3a sycl : Add Q5_K ESIMD kernel (llama/26376) df8336ef opencl: keep the vocab-scale K-quant lm_head on the CPU for Adreno A7X (compiler issue workaround) (llama/26440) 99694692 sycl: Update gate logic for Alchemist GPUs regarding OneDNN features. (llama/26635) 1d30b1b5 sycl: fix multiple warnings in compiling sycl backend (llama/26713) 7378ac17 sycl : fix load model with mlock issue (llama/27250) 4979ee24 ggml: support ggml_rope_set_offset on opencl, sycl, wgpu, hexagon (llama/27345) 33c9ea5e metal : clamp K extent in tensor API mat-mat kernel for K not a multiple of 32 (llama/27450) 9e7a4c2c opencl: fix q6_K flat mul_mat for Adreno A6x/A7x GPUs with older E031 compilers (llama/26476) 505842b4 opencl: fix local size for norm (llama/27339) b0d45de2 vulkan: FA MMQ should use fp32 for Q quantization calculations (llama/27413) 3821f6ed metal : dequant kv cache only for large batches (llama/27438) 7159fc6a CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows (llama/26678) 02a0ab2b CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover (llama/26079) e4159ddd metal : dequantize quantized KV to F16 before flash attention (llama/27390) 9a1d2348 Revert "tensor-split meta backend fixes (#26502)" (llama/27433) 0ab579a5 ggml: fix backend split scheduler race condition (llama/26040) ee401c8b ggml-cuda: provide static workspace for cuBLAS handles (llama/26574) 6a6a4c19 vulkan : add source groups for shaders (llama/26666) e3e8dc16 opencl: make the MoE expert scatter deterministic (llama/26464) 7c6058cf tensor-split meta backend fixes (llama/26502) 37c79ff5 hexagon: fix FA HMX queue ordering and pack the rescale D matrices (llama/27042) 585e463c opencl: port fused ssm_scan kernel (Mamba-2, d_state in {128, 256}) to GPU (llama/26439) ff4f3e43 ggml-cpu: gate __fp16 on __ARM_FP16_FORMAT_IEEE (llama/26860) 480cc77c vulkan : dequant q8_0 KV once in coopmat1 (llama/25494) 8fc9724d vulkan: add null checks in ggml_vk_queue_command_pools_cleanup (llama/27353) f673848d sycl: report zero devices instead of aborting when the host has none (llama/27291) aa9d2226 ggml: add ggml_rope_set_offset (+ metal support) (llama/27120) 868403cf metal : dequantize q8_0 using packed types (llama/27370) 55e4b3b4 vulkan: tiled transpose for 0<->2 permuted CONT (llama/26585) 52b66a4e ggml-webgpu: add mulmat with overlapping src0/src1 (e.g., for minimax-01) (llama/27321) 981a41bb opencl: fix WAR race in the generic FA tile kernels when the WG spans subgroups (llama/26434) f16a1a2c RPC: populate use_count to enable fusion inside backends (llama/27142) b74262dd sycl: honor GGML_HINT_SRC0_IS_HADAMARD (llama/27298)

Source: README.md, updated 2026-08-21