Download Latest Version v1.9.5 source code.zip (10.7 MB) Google Add to Preferred Sources
Home / v1.9.5
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-10-06 24.3 kB
v1.9.5 source code.tar.gz 2026-10-06 9.6 MB
v1.9.5 source code.zip 2026-10-06 10.7 MB
Totals: 3 Items   20.3 MB 105

Overview

New version has been released.

Nightly build: b5454 More info: dist : releases and versioning of ggml-org projects

Changelog since v1.9.4

d1be6fde whisper : bump version to 1.9.5 (#4102) 4afec37b talk-llama : update llama.cpp to v0.6.0 3d451abe sync : ggml 47f1c7dc ggml : bump version to 0.26.0 (ggml/1652) e71b7844 CUDA: make the alloc_deps check batch independent (llama/29986) cdc87302 vulkan: fix Flash Attention shmem write out of bounds (llama/29988) 563a9f7a vulkan: revert mul_mat_id tile selection PR [#29182] (llama/29936) ebb8a170 cuda: use the vector lightning indexer kernel on MUSA (llama/29990) 8c6ca4ab CUDA: Optimize accumulation in mmq for NVFP4 type (llama/29857) 82fa9df7 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (llama/29591) f6915422 vulkan: sparse flash attention for quantized K/V (llama/29639) d477de3c llama : fix unexpected graph reallocation in the k-pool models (llama/29958) 95283ca7 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (llama/28479) 155c5c28 vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (llama/29912) 7091d4a2 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (llama/29483) 4cc62b81 cuda: tile the lightning indexer over keys and tokens for 4 heads (llama/29901) 7e576cbd metal : few-row MMA mat-mul (llama/29869) 91da4702 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (llama/29633) 8bb85d96 CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (llama/29435) ce6552ef CUDA: refactor swizzling code (llama/29612) bfe3dc63 ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (llama/29806) 25453c0a cuda : move neu_padded to where it is used (llama/29940) 7914776a cuda : move blocks_per_col to where it is used (llama/29939) 6421d2c1 CUDA: fix MMQ memory fault if n_expert >> n_ubatch (llama/29941) 05d4d89a vulkan: fix rdna4 mat_vec tuning (llama/29934) 51bfbd61 webgpu: add f16 support to fill/set_rows (llama/29897) f2aa80ad ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (llama/29852) 7f4a67a5 qwen4exp : halve the indexer score memory (llama/29825) cf5d9e30 CUDA: fuse shared experts into MMVQ (llama/29184) d7682b58 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (llama/27663) 6b704713 ggml-quants : avoid invalid rounding in qkx3 scale search (llama/29817) 0b35d1f3 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (llama/27096) d0dc4363 metal : add tensor API flash attention kernel for F16 KV (llama/29570) cca8f72b opencl: use sigmoid f16 for bf16 (llama/29787) cd1bee52 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (llama/29186) 95a05ed4 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (llama/28531) c4051a33 sycl: large register file for D=512 FA vec kernels (llama/29062) ae92605f sycl : do not use slow oneDNN reference matmul and fattn (llama/28985) 9c968f7e qwen4exp : optimize mask constructions (llama/29824) 4435763c ggml : add alloc_buffer_n to buffer type interface (llama/23671) 8291ab84 vulkan: add logging to pipeline compile issues (llama/29794) eaadab3d hexagon: install rebuilt HTP skels (llama/29828) 0295ef68 hexagon: add q2_k and q3_k quant type support (llama/29717) 396f68d3 CUDA: fix 2 broken Volta FA cases (llama/29803) 3a3598cc llama: refer to segment documentation [no ci] (llama/29074) 15229ac9 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (llama/29685) aa538005 cuda : route sm70 to the Turing MMVQ nwarps table (llama/29753) 81ca4f81 metal : release temporary private transfer buffers (llama/29777) 88948bdb webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- [#29358] (llama/29358) 10872be7 CUDA: Handle compute type for NVFP4 on cublass path (llama/29173) 3b68f901 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (llama/29792) 56500f4f meta: clear inactive AllReduce shards with FILL, not SCALE (llama/29793) 817294a9 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (llama/29572) b4a085a0 hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (llama/29785) 0ecc317e BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (llama/29640) 67ad86d0 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (llama/29698) 61ea5027 metal : use bf16 math for mxfp4 mul-mat (llama/29770) d371e376 webgpu: fix SSM_SCAN binding aliasing (llama/29750) c1c13908 ggml-opencl : replace alloca() with std::vector (llama/29765) bffb6e4b cuda: guard the iq4_nl dequantize row kernel against short rows (llama/29683) 6773ef87 Hexagon: optimize ALLREDUCE with support for safe scatter mode (llama/29757) 02ea0c2e ggml/gguf : fix integer overflow (llama/29384) 067bb06b ggml-et : remove useless alloca() (llama/29663) fd66b6ec ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (llama/29675) c1b3fb10 cpu: accept BF16 in src1 of mul_mat (llama/28937) d0af734d openvino: serve GET_ROWS on a weight view from the base Constant (llama/28381) 1cbf7e79 musa : define CUDA_ARCH for device passes (llama/29508) d03bfd7f SYCL: reduce tensor allreduce sync with pinned host buffers (llama/29604) 66e8b957 ggml-zdnn: impl buffer reset, fix memory leaks (llama/29637) b8051db8 Hexagon f16 activation ops (llama/29209) 6e48a365 gguf : reject tensor size that wraps after padding (llama/26979) 14423995 ggml : check row bounds in get_rows_back (llama/29575) 8d2a2cb4 hexagon: optimize concat op (llama/29673) 45093fc8 CUDA: bitonic argsort handles rows wider than one block (llama/28957) 482baacd ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (llama/29478) 3b86cabd ci: add zdnn backend build but not test (llama/29541) 9f12c9fe opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (llama/29555) 5439f3d9 vulkan: Tune GDN kernel, fix Intel performance (llama/29476) 6c07d821 vulkan : Load F32 A matrix 2 at a time when its 2-aligned (llama/29254) 8cc3ca48 vulkan: MOE aware mat_mul_id tile selection (llama/29182) deae6f34 ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (llama/29504) b97d4193 ggml : accumulate f16 dot products in f32 on AVX512-FP16 (llama/29545) 4d4817ed ggml : require input tensors to be GGML_OP_NONE (llama/29647) 5e1914a7 hexagon: add FP32 GELU_ERF and GEGLU_ERF support (llama/29631) 763b67fd ggml : collect all input tensors into graph_inputs (llama/29634) 84781290 ggml-zdnn: fix 0-row tensor crash (llama/29636) f449de94 ggml : speed up model loading (llama/29598) b30ef05f metal: FWHT perf optimizations (llama/29602) eecee67b vulkan : reuse descriptor sets when bindings are constant (llama/29280) 97690105 vulkan: include functional header (llama/29597) 830dc664 ggml-openvino: mark unaligned batch-stride views unsupported (llama/29603) 81b951e0 webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (llama/29471) ee86a150 tests : refactor test-recurrent-state-rollback (llama/29426) 83fd1554 ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (llama/29423) 65449dbc vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (llama/29520) 0c17f7d6 HIP: fix template skip for DKQ > 256 mfma kernels (llama/29559) 0e537510 metal: support left and circular padding in GGML_OP_PAD (llama/29561) b7b84950 Enables Windows ARM64 build with MSVC cl.exe (llama/28362) 85f69261 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (llama/28956) 51db4755 opencl: refine bin kernel loading condition (llama/29503) 55a695da sycl: FWHT kernels for block widths above 512 (llama/29243) 846af526 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (llama/26289) 743f1ad2 HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (llama/28907) 14d1aa70 vulkan: fix argsort kernel selection for Adreno (llama/29469) 24cf265b RPC: use RDMA completion channel to not spin (llama/29440) 7eea5188 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (llama/29511) fe06027d hexagon: support for backend sampler (llama/29502) fc1ebfab cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (llama/28717) dfe8fbd7 cuda: add F16 input to the FWHT (llama/29096) 447a7751 ggml-cpu: tiled mul_mat for k-quants (llama/27851) bf1d787a opencl: add A8 Q8_0 non-MoE dp4a binary kernel (llama/29439) 9731f6fd hexagon: find software divide calls using binary inspection tool (llama/29449) 2d91497a opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (llama/29401) 669188ef Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (https://github.com/ggml-org/llama.cpp/issues/29373) (llama/29409) 9a919424 metal: FWHT kernels for block widths above 512 (llama/29095) 331d3c8a sync : ggml 4b0998ad metal : split fa kernels into per-dtype libraries (llama/29329) bad9b546 sync : ggml afc10c26 llama : add llama_prec_policy + model-driven W4A4 path (llama/24364) d7e83b18 HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (llama/29231) 26bfdc66 rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (llama/29283) 93c4308a support sparse FA (llama/28796) f7d5bf38 musa: fix PH1 (MTT S5000) operator failures and build issues (llama/29193) 40c872b2 CUDA: fuse RMS_NORM + SCALE into one kernel (llama/29393) 0458c9eb hexagon: add q5_k quant type support (llama/29123) eea9aa63 hexagon: use DMA for contiguous dim1 CONCAT (llama/29404) 4550b9db metal : fix graph capture and handle empty graphs (llama/29390) e1b93a51 metal : optimize sparse FA + clean-up (llama/29377) 61032a2b hexagon: handle multi-sequence in concat_2d (llama/29344) 845d1c02 hexagon: dynamic quantizer improvements (llama/29395) e5a88776 hexagon: support I32 CPY and CONT (llama/29379) 5a9a3c06 cuda : add F16 kernel support for CONV_2D_DW (llama/29064) cf0852d2 vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (llama/27952) 6d0197cc ggml : bump version to 0.25.3 (ggml/1645) a2b9a22e ggml : fix ubsan error in ggml_graph_nbytes (ggml/1644) 7ab1529b ggml : bump version to 0.25.2 (ggml/1642) f938f319 vulkan: handle misalignment in conv_2d and conv_3d (llama/29365) bf8128ec vulkan: tune KHR cooperative matrix support for Adreno GPUs (llama/29328) 3ef6a832 cuda : add conv3d with implicit GEMM (llama/29137) 6f2a84af hexagon: reject MUL_MAT_ID when src1 precision is F32 (llama/29348) 0adffd58 opencl: add A8 Q6_K non-MoE dp4a binary kernel (llama/29057) 60c0be6a whisper : add check for ttype in whisper and parakeet (#4091) 6e4ab854 cli : fix incorrect error code when files are failing (#4080) d09f61a7 ci : cover GGML_BACKEND_DL in ubuntu-22-clang-arm64 (#4047) a664346e sync : ggml 84f080e9 ggml : bump version to 0.25.1 (ggml/1637) 8f5ac2a6 CUDA: add a reserve to avoid spurious warning on older GCC builds (llama/29317) 8917ea04 metal: add the missing f32 x bf16 mul_mv variants (llama/28741) f48aebe2 CUDA: enable sparse-fa for dsv4 prefill (again) (llama/29298) 6bc51cc7 metal : key the fa-vec tuned table by family instead of SKU (llama/29075) 18521323 vulkan: add IQ4_XS MMQ/MMV matmul kernels (llama/28415) ed1339a5 sync : ggml 3e7723fe ggml : bump version to 0.25.0 (ggml/1635) 711ef844 common : fix for two functions when top_k exceeds the vocabulary size. (ggml/1633) f6b039ff sycl : fix compile warnings 431ecf57 ggml-meta: resolve multi buffer views (llama/29266) e0ca36d0 cuda: top-k MoE should always fire (llama/28432) d6075ffc sycl : support new UT case for mul_mat_hadamard fp16 (llama/29218) 6570b79e sycl: extend MMVQ GLU fusion, add rms_norm+scale and ssm_conv+silu fusions (llama/28931) a1c7f973 sycl : support op get_rows_back, only support fp32/fp16 (llama/25266) fbcb94d0 vulkan: hide internal symbols to prevent duplicate-dlopen state destruction (llama/29139) b18bad0d hex-dma: introduce direct-mapped DMA cache that is better suited for HVX FA mask handling (llama/29282) e59366ba HIP : optimize IQ2/IQ3 (__vsub4 __vcmpne4) using SWAR (llama/27962) 736efcf9 opencl: add bin kernel kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin (llama/29056) e900c888 vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LPG Plus/Xe2/Xe3) (llama/24406) ad0058be metal : gate mul_mm_id src1 rescale behind ggml_prec (llama/29029) ff565cf3 ggml : IQ1_M build prefix sums once per block (llama/28706) 0409ed26 Performance tune for gemma4-26b-a4b flash attention shape. (llama/28450) dd67678e opencl: add A8 Q4_0 non-MoE dp4a binary kernel (llama/29055) 898392b7 hexagon: new HMX-optimized GATED_DELTA_NET (llama/29199) 63412d34 metal : fix mask bounds in flash attention block pre-pass (llama/29220) 0e640a18 cuda: fix sm_70 tile compilation error (llama/29224) 9a7d43df ggml-cuda : convert contiguous tensors four elements at a time (llama/29155) 86ef1b15 cuda : accelerate conv2d with implicit GEMM (llama/29135) d18a4222 ggml : fix dimension and stride truncation in ggml_permute (llama/29227) eb279cf6 sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix (llama/29132) 0ecf57d7 sycl : pinned memory use right device context instead of 0 (llama/28895) 34a4c124 CUDA: Follow up of [#25635], refactoring FA shared smem swizzle (llama/28536) b7b1fe47 ggml-metal : simplify fusion pattern op list declaration (llama/29206) 42c873a2 sycl : coalesce MKL-FA softmax loads instead of one work-item per row (llama/28918) b0c51495 ggml-cpu: ARM Repack kernels for Q1_0 (llama/23492) b50d5c31 hexagon: overhaul of buffer and DMA handling to support 64bit mappings + improvements (llama/29197) 76d02a12 cuda : tune MMVQ to MMQ crossover for SM70 (Volta) (llama/28912) 87b6876b metal : fix deprecation warnings from macOS 27 SDK (llama/29136) 29e710c4 webgpu : add fused gdn + cpy (llama/28976) 4ef9fd70 CUDA: tune FA for Gemma 4 on Ampere or newer (llama/29152) 984e400c metal : support arbitrary hc in dsv4_hc_pre (llama/29169) 3d949a36 CUDA: enable sparse fa for qwen4 (llama/28770) c77f6bbe metal: add F16 input to the FWHT (llama/29094) 7f5ac73e hexagon: enable I32 GET_ROWS (llama/29116) 232d7189 hexagon: add support for GEGLU_QUICK (llama/29114) 099090cf hexagon: enable support for TOP_K op (llama/29113) a00fa32e metal : add MoE and SSM_CONV fusion optimizations (llama/28948) 47456f6f metal : fix FA support checks (llama/29122) 4f1cce14 metal : support qwen4exp hc ops (llama/29000) 58844b02 cuda : fix CUB argsort corruption caused by in-place keys (llama/28389) 97dc0171 opencl: add support for bin kernel flash_attn_f32_f16_bin (llama/29046) b29b4398 hexagon: add ROLL op support (llama/29105) e2358dff hexagon: im2col update (llama/29103) c01abcec hexagon: HMX flash-attention head_dim padding (support DK=DV=72) (llama/26539) 1b6c6a95 opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin (llama/28678) 8f386039 ggml-cpu: add F16 input to the FWHT (llama/27779) bdf289e0 ggml-webgpu: fix supports_op condition for GET_ROWS (llama/28978) 26d6dcfc ggml : handle graph buffer reservation failure (llama/26070) 8e336cd0 vulkan: add IQ3_S MMQ matmul kernels (llama/28822) 6b752e5e vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (llama/28501) f455712a openvino : Update OpenVINO to 2026.4;fix clangd,MSVC warnings; (llama/29009) fbdbbc78 vulkan: split buffers and debug code into separate files, add shared headers (llama/28732) 78264220 gguf : align the data section relative to the GGUF start, not the file (llama/28993) b9e5f3ac sycl : fix the B70 mem allocate error when >19.3GB (llama/28953) 84a020e3 vulkan: skip unneeded MoE work in mul_mm coopmat1 path (llama/25483) cb44896d sycl: ssm_conv: fuse the SiLU epilogue into the ssm_conv kernel (llama/28929) d4b9101e opencl: fix various warnings (llama/28984) d375e3c1 vulkan: fix buffer_reference alignment in im2col shaders (llama/28996) 31c972a9 vulkan: support qwen4exp hc ops (llama/28988) d673fbac Fix function signature for ggml_backend_sycl_split_buffer_type (llama/28981) 67cf515a vulkan: work around NV bug with argsort_large.comp (llama/28975) 6e220ab6 hexagon: Support for K-Quants Q4_K and Q6_K (llama/28994) 6ca20bb3 hexagon: accept the zeroed rope probe in supports_op (llama/28995) 94d6e24e CUDA/HIP: improve access patterns in im2col (llama/28013) af5153ce spacemit : fix wrong transpose function for int16 data (llama/25161) dce52b19 rpc : invalidate cached compute graph when a referenced buffer is freed (llama/24292) fa2c801c qwen4exp: add hc ops (llama/28901) f9ad9868 HIP: broaden MoE ncols_opt tile heuristic on RDNA3.5 architecture (llama/28935) 0b9fb0f2 vulkan: make MUL_MAT_ID BN/2 tail unconditional (llama/28923) edbb13e4 metal: fix NaN in mul_mm_id when activations exceed f16 range (llama/26223) 7dd0ce90 hexagon: add back missing contiguous fast-path and hvx_copy_uu for each run (llama/28886) 9e6e308e hex-cpy: use dma if src and dst are contiguous (llama/28906) 43f7501f HIP: Enable AllReduce for ROCm (llama/27825) bae0f977 opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (llama/27637) 1fc756a2 rpc : hash-cache only weights (llama/28789) 6189e6fc cuda: support row-contiguous SUM_ROWS (llama/26308) f69f5904 vulkan: support sparse Flash Attention (llama/28105) 4c341e28 OpenVINO: optimize stateful decode and GPU MoE inference (llama/28638) 67630d00 opencl: add generic ssm_scan (llama/28881) 4be99fe7 metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (llama/28599) b1fd0cab cuda : enable i16 and i32 for DUP (llama/28897) 50e4eedd HIP: fattn-mma: use fp32 accumulation on MFMA devices (llama/28576) 398997ed whisper : add abort_callback on lang detection (#4077) a44e0784 vad : reject n_encoder_layers other than 4 in model load (#4064) 307869af devops : reduce Vulkan Docker image to 39.4% of its original size (now 692 MB) (#4038) 5670d5c0 fix(yt-wsp): Resolve script path without GNU realpath (#4072) b27fbff4 cli : load backends after validating input files (#4069) fd7d8abb ci : update android-actions to v4.0.4 (#4074) d5d6e59b docs : clarify VAD mode timestamps and CWD model path errors (#4019) 7a2ceef9 readme : document the ANEForge encoder backend (#4073) 4afa0094 whisper : optional ANEForge encoder backend (Apple Neural Engine) (#3905) da545722 whisper : fix int overflow in whisper_full_parallel chunk offsets (#4044) 1d549b3c sync : ggml ce5c557b ggml : bump version to 0.24.0 (ggml/1627) 950a4a42 tests(s390x): add non-vxe build to tests (llama/28776) 0a866737 sycl: rfc: Use radix select for top_k (llama/28670) a3d26001 ggml-cpu : disable PCH and fix CACHE_LINE_SIZE ambiguity to fix heap corruption (llama/28882) 0bea2893 sycl : fix oneDNN scratchpad breaking the pool free order (llama/28704) ac78ae98 ggml-cuda: fallback to F32 on device without BF16 hardware acceleration (llama/28846) 8f4cd19c ggml-cpu(s390x): guard VXE-only repack helpers (llama/28775) 51ee9279 sycl : Fix get mem error (llama/28227) f59047cf vulkan: workaround NV queuesubmit driver bug (llama/28830) 17e63925 opencl: apply the noshuffle row-alignment rule to q4_K, q5_K and q8_0, not just q6_K (llama/28575) 8751eaeb ggml-cuda: hip add specific config table for AMD GCN (llama/27841) 24035835 syscl : Handle (fail gracefully) unsupported tq1_0 quants (llama/28681) df31856c rpc : fix linking when compiling with BUILD_SHARED_LIBS=OFF (llama/28492) 07825dc1 opencl: fix several bugs where the backend aborts (llama/27630) ec82a96f opencl: add bin kernel kernel_gemm_noshuffle_q4_k_f32_32b_trans_ila_a8_bin (llama/28677) b0e076d6 webgpu: align tensor bindings to the type block size (llama/28382) ea3ef8f5 hexagon: support for multi-device model split (aka row-split) (llama/28589) b1779413 ggml-webgpu: Update to a recent version of Dawn (llama/28683) 2c1b5257 ggml: skip 0-sized ids tensor when offloading selected experts (llama/28739) 5f310fa1 metal : skip the empty half of the mul_mm_id token tile (llama/28301) 76da3529 cmake : add PCH and unity build to improve build times (llama/28091) 9d70ac2f metal : single-source fusion table + fusion debug rework (llama/28164) 59cca2c7 metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 (llama/28692) 61b318f4 CUDA/HIP: Flash Attention tuning (gfx1201) (llama/28102) 468710cc vulkan: fix data race and OOB access in argsort(large) (llama/28705) e5369a69 opencl: add A8 Q4_0 mm binary kernel support (llama/28268) 14e868eb vulkan: use CPU writes in ggml_backend_vk_cpy_tensor_async if the context is idle (llama/28618) fb8f4274 vulkan: use add_alloc_dep to enable topk_moe fusion for prefill (llama/28422) e955658b vulkan: small M matrix optimizations for qwen (llama/28457) d374147d vulkan: fall back to shared-memory reduction for dmmv on PowerVR (llama/28341) a6e85dd5 vulkan : add command-buffer debug labels for GPU profilers (llama/28101) 9f7331f7 ggml-cpu(s390x): add repack support for q4_0 (llama/28667) d3945bf9 ggml-cpu(s390x): add Q1_0 vector intrinsic support (llama/28606) 3a54d53e vulkan: use spec constant for matrix matrix multiplication A-type (llama/25773) bafeaca8 hexagon: rope updates (llama/28628) 84b8db90 vulkan: Convert FILL to distribute workgroups in 2D to avoid exceeding maxComputeWorkGroupCount (llama/28592) f55b67f0 CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (llama/28552) 4d506f58 CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (llama/28079) dca2df42 vulkan: add dedicated iq4_xs mat-vec shader (llama/28426) 355d90b4 vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 (llama/27471) facf4b50 Add IQ type handling for MoE (llama/28476) 32a749dc ggml : fix msvc+clang ggml_vld1q_u32 (llama/28284) 006e53e4 Revert "ggml-cuda : restore prop.integrated on HIP builds (llama/24233)" (llama/28604) 5b984949 metal : fix idle threads in mul_mv_iq3_xxs for ne00 < 1024 (llama/28086) 8bae0820 llama : add missing headers (llama/28566) b1275431 vulkan : fuse UNARY(GELU|SIGMOID|SILU|SOFTPLUS) + MUL (llama/27220) 69fcec3b Fix Vulkan-Hpp handle usage on 32-bit targets. (llama/22892) 37b210d8 opencl: properly handle non-contiguous inputs to conv2d (llama/28503) 33cadde1 ggml : update ggml_prec specification (llama/26675) 707c3ee1 hexagon: add RELU and LEAKY_RELU ops (llama/28585) 34f53363 Revert "CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (#24546)" (llama/28551) 0f9591a1 webgpu: format the GET_ROWS case block (llama/28542) b61186de sycl: add a batched L2_NORM kernel (llama/28222) 8ce432f1 vulkan: add DeepSeek-V4 hyper-connection fused ops (DSV4_HC_COMB/PRE/POST) (llama/26578) 4c5a9a02 CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (llama/24546) f15e1a67 ggml: add gfx90c HIP support (llama/26454) a11d16a8 CUDA: branchless Q4_K/Q5_K unpack to speed up mmvq, L2 prefetch on DGX Spark (llama/26705) 60e475f2 vulkan: support type-aligned GET_ROWS (llama/28253) c6135ac2 ggml-cuda: fix divergent barrier in f16 flash attention (llama/27870) 1da558c5 ggml: allow backend inputs to not create another split (llama/28387) 8d3ed200 vulkan: rms_norm fusion opportunities (llama/28024) 088b3b4a vulkan: add TQ1_0 support (mm, mat-vec, mat-vec-id, dequant, get_rows) (llama/27765) ad344d16 opencl: properly choose weights pack for q4_K, q5_K mul_mat (llama/28402) a0e75c84 cuda: fixes races in mmid and mmf (llama/28475) 1299eb6c metal : add remaining fa-vec tunings for M2 Max (llama/28458) 614c74ab metal : fix memory leak in early return (llama/28399) 0006e3e2 sycl : fix test-backend-ops CI break && restore Kronecker product FWHT support (#28016) (llama/28254) 07796107 sycl: attribute device allocations by site (GGML_SYCL_MEMTRACE) (llama/27631) ad218df4 metal : add remaining fa-vec tunings for M3 (llama/28396) 25c6ab5d opencl: extend the elementwise and data‐movement op coverage (#27633) 7fb2dafc opencl: add Adreno xmem SDPA path (llama/26331) 70598eee scripts : fix sync (#0) f133970b ci : use devlab-dispatch for npu-amd-windows (#4060) 1da4dc82 ci : rename cublas to cuda in release.yml (#4057) 02612981 ci : use devlab-dispatch for npu-amd-linux (#4056)

Source: README.md, updated 2026-10-06