| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-10-06 | 24.3 kB | |
| v1.9.5 source code.tar.gz | 2026-10-06 | 9.6 MB | |
| v1.9.5 source code.zip | 2026-10-06 | 10.7 MB | |
| Totals: 3 Items | 20.3 MB | 105 | |
Overview
New version has been released.
Nightly build: b5454 More info: dist : releases and versioning of ggml-org projects
Changelog since v1.9.4
d1be6fde whisper : bump version to 1.9.5 (#4102)
4afec37b talk-llama : update llama.cpp to v0.6.0
3d451abe sync : ggml
47f1c7dc ggml : bump version to 0.26.0 (ggml/1652)
e71b7844 CUDA: make the alloc_deps check batch independent (llama/29986)
cdc87302 vulkan: fix Flash Attention shmem write out of bounds (llama/29988)
563a9f7a vulkan: revert mul_mat_id tile selection PR [#29182] (llama/29936)
ebb8a170 cuda: use the vector lightning indexer kernel on MUSA (llama/29990)
8c6ca4ab CUDA: Optimize accumulation in mmq for NVFP4 type (llama/29857)
82fa9df7 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (llama/29591)
f6915422 vulkan: sparse flash attention for quantized K/V (llama/29639)
d477de3c llama : fix unexpected graph reallocation in the k-pool models (llama/29958)
95283ca7 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (llama/28479)
155c5c28 vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (llama/29912)
7091d4a2 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (llama/29483)
4cc62b81 cuda: tile the lightning indexer over keys and tokens for 4 heads (llama/29901)
7e576cbd metal : few-row MMA mat-mul (llama/29869)
91da4702 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (llama/29633)
8bb85d96 CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (llama/29435)
ce6552ef CUDA: refactor swizzling code (llama/29612)
bfe3dc63 ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (llama/29806)
25453c0a cuda : move neu_padded to where it is used (llama/29940)
7914776a cuda : move blocks_per_col to where it is used (llama/29939)
6421d2c1 CUDA: fix MMQ memory fault if n_expert >> n_ubatch (llama/29941)
05d4d89a vulkan: fix rdna4 mat_vec tuning (llama/29934)
51bfbd61 webgpu: add f16 support to fill/set_rows (llama/29897)
f2aa80ad ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (llama/29852)
7f4a67a5 qwen4exp : halve the indexer score memory (llama/29825)
cf5d9e30 CUDA: fuse shared experts into MMVQ (llama/29184)
d7682b58 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (llama/27663)
6b704713 ggml-quants : avoid invalid rounding in qkx3 scale search (llama/29817)
0b35d1f3 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (llama/27096)
d0dc4363 metal : add tensor API flash attention kernel for F16 KV (llama/29570)
cca8f72b opencl: use sigmoid f16 for bf16 (llama/29787)
cd1bee52 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (llama/29186)
95a05ed4 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (llama/28531)
c4051a33 sycl: large register file for D=512 FA vec kernels (llama/29062)
ae92605f sycl : do not use slow oneDNN reference matmul and fattn (llama/28985)
9c968f7e qwen4exp : optimize mask constructions (llama/29824)
4435763c ggml : add alloc_buffer_n to buffer type interface (llama/23671)
8291ab84 vulkan: add logging to pipeline compile issues (llama/29794)
eaadab3d hexagon: install rebuilt HTP skels (llama/29828)
0295ef68 hexagon: add q2_k and q3_k quant type support (llama/29717)
396f68d3 CUDA: fix 2 broken Volta FA cases (llama/29803)
3a3598cc llama: refer to segment documentation [no ci] (llama/29074)
15229ac9 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (llama/29685)
aa538005 cuda : route sm70 to the Turing MMVQ nwarps table (llama/29753)
81ca4f81 metal : release temporary private transfer buffers (llama/29777)
88948bdb webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- [#29358] (llama/29358)
10872be7 CUDA: Handle compute type for NVFP4 on cublass path (llama/29173)
3b68f901 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (llama/29792)
56500f4f meta: clear inactive AllReduce shards with FILL, not SCALE (llama/29793)
817294a9 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (llama/29572)
b4a085a0 hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (llama/29785)
0ecc317e BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (llama/29640)
67ad86d0 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (llama/29698)
61ea5027 metal : use bf16 math for mxfp4 mul-mat (llama/29770)
d371e376 webgpu: fix SSM_SCAN binding aliasing (llama/29750)
c1c13908 ggml-opencl : replace alloca() with std::vector (llama/29765)
bffb6e4b cuda: guard the iq4_nl dequantize row kernel against short rows (llama/29683)
6773ef87 Hexagon: optimize ALLREDUCE with support for safe scatter mode (llama/29757)
02ea0c2e ggml/gguf : fix integer overflow (llama/29384)
067bb06b ggml-et : remove useless alloca() (llama/29663)
fd66b6ec ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (llama/29675)
c1b3fb10 cpu: accept BF16 in src1 of mul_mat (llama/28937)
d0af734d openvino: serve GET_ROWS on a weight view from the base Constant (llama/28381)
1cbf7e79 musa : define CUDA_ARCH for device passes (llama/29508)
d03bfd7f SYCL: reduce tensor allreduce sync with pinned host buffers (llama/29604)
66e8b957 ggml-zdnn: impl buffer reset, fix memory leaks (llama/29637)
b8051db8 Hexagon f16 activation ops (llama/29209)
6e48a365 gguf : reject tensor size that wraps after padding (llama/26979)
14423995 ggml : check row bounds in get_rows_back (llama/29575)
8d2a2cb4 hexagon: optimize concat op (llama/29673)
45093fc8 CUDA: bitonic argsort handles rows wider than one block (llama/28957)
482baacd ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (llama/29478)
3b86cabd ci: add zdnn backend build but not test (llama/29541)
9f12c9fe opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (llama/29555)
5439f3d9 vulkan: Tune GDN kernel, fix Intel performance (llama/29476)
6c07d821 vulkan : Load F32 A matrix 2 at a time when its 2-aligned (llama/29254)
8cc3ca48 vulkan: MOE aware mat_mul_id tile selection (llama/29182)
deae6f34 ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (llama/29504)
b97d4193 ggml : accumulate f16 dot products in f32 on AVX512-FP16 (llama/29545)
4d4817ed ggml : require input tensors to be GGML_OP_NONE (llama/29647)
5e1914a7 hexagon: add FP32 GELU_ERF and GEGLU_ERF support (llama/29631)
763b67fd ggml : collect all input tensors into graph_inputs (llama/29634)
84781290 ggml-zdnn: fix 0-row tensor crash (llama/29636)
f449de94 ggml : speed up model loading (llama/29598)
b30ef05f metal: FWHT perf optimizations (llama/29602)
eecee67b vulkan : reuse descriptor sets when bindings are constant (llama/29280)
97690105 vulkan: include functional header (llama/29597)
830dc664 ggml-openvino: mark unaligned batch-stride views unsupported (llama/29603)
81b951e0 webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (llama/29471)
ee86a150 tests : refactor test-recurrent-state-rollback (llama/29426)
83fd1554 ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (llama/29423)
65449dbc vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (llama/29520)
0c17f7d6 HIP: fix template skip for DKQ > 256 mfma kernels (llama/29559)
0e537510 metal: support left and circular padding in GGML_OP_PAD (llama/29561)
b7b84950 Enables Windows ARM64 build with MSVC cl.exe (llama/28362)
85f69261 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (llama/28956)
51db4755 opencl: refine bin kernel loading condition (llama/29503)
55a695da sycl: FWHT kernels for block widths above 512 (llama/29243)
846af526 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (llama/26289)
743f1ad2 HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (llama/28907)
14d1aa70 vulkan: fix argsort kernel selection for Adreno (llama/29469)
24cf265b RPC: use RDMA completion channel to not spin (llama/29440)
7eea5188 hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (llama/29511)
fe06027d hexagon: support for backend sampler (llama/29502)
fc1ebfab cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (llama/28717)
dfe8fbd7 cuda: add F16 input to the FWHT (llama/29096)
447a7751 ggml-cpu: tiled mul_mat for k-quants (llama/27851)
bf1d787a opencl: add A8 Q8_0 non-MoE dp4a binary kernel (llama/29439)
9731f6fd hexagon: find software divide calls using binary inspection tool (llama/29449)
2d91497a opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (llama/29401)
669188ef Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (https://github.com/ggml-org/llama.cpp/issues/29373) (llama/29409)
9a919424 metal: FWHT kernels for block widths above 512 (llama/29095)
331d3c8a sync : ggml
4b0998ad metal : split fa kernels into per-dtype libraries (llama/29329)
bad9b546 sync : ggml
afc10c26 llama : add llama_prec_policy + model-driven W4A4 path (llama/24364)
d7e83b18 HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (llama/29231)
26bfdc66 rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (llama/29283)
93c4308a support sparse FA (llama/28796)
f7d5bf38 musa: fix PH1 (MTT S5000) operator failures and build issues (llama/29193)
40c872b2 CUDA: fuse RMS_NORM + SCALE into one kernel (llama/29393)
0458c9eb hexagon: add q5_k quant type support (llama/29123)
eea9aa63 hexagon: use DMA for contiguous dim1 CONCAT (llama/29404)
4550b9db metal : fix graph capture and handle empty graphs (llama/29390)
e1b93a51 metal : optimize sparse FA + clean-up (llama/29377)
61032a2b hexagon: handle multi-sequence in concat_2d (llama/29344)
845d1c02 hexagon: dynamic quantizer improvements (llama/29395)
e5a88776 hexagon: support I32 CPY and CONT (llama/29379)
5a9a3c06 cuda : add F16 kernel support for CONV_2D_DW (llama/29064)
cf0852d2 vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (llama/27952)
6d0197cc ggml : bump version to 0.25.3 (ggml/1645)
a2b9a22e ggml : fix ubsan error in ggml_graph_nbytes (ggml/1644)
7ab1529b ggml : bump version to 0.25.2 (ggml/1642)
f938f319 vulkan: handle misalignment in conv_2d and conv_3d (llama/29365)
bf8128ec vulkan: tune KHR cooperative matrix support for Adreno GPUs (llama/29328)
3ef6a832 cuda : add conv3d with implicit GEMM (llama/29137)
6f2a84af hexagon: reject MUL_MAT_ID when src1 precision is F32 (llama/29348)
0adffd58 opencl: add A8 Q6_K non-MoE dp4a binary kernel (llama/29057)
60c0be6a whisper : add check for ttype in whisper and parakeet (#4091)
6e4ab854 cli : fix incorrect error code when files are failing (#4080)
d09f61a7 ci : cover GGML_BACKEND_DL in ubuntu-22-clang-arm64 (#4047)
a664346e sync : ggml
84f080e9 ggml : bump version to 0.25.1 (ggml/1637)
8f5ac2a6 CUDA: add a reserve to avoid spurious warning on older GCC builds (llama/29317)
8917ea04 metal: add the missing f32 x bf16 mul_mv variants (llama/28741)
f48aebe2 CUDA: enable sparse-fa for dsv4 prefill (again) (llama/29298)
6bc51cc7 metal : key the fa-vec tuned table by family instead of SKU (llama/29075)
18521323 vulkan: add IQ4_XS MMQ/MMV matmul kernels (llama/28415)
ed1339a5 sync : ggml
3e7723fe ggml : bump version to 0.25.0 (ggml/1635)
711ef844 common : fix for two functions when top_k exceeds the vocabulary size. (ggml/1633)
f6b039ff sycl : fix compile warnings
431ecf57 ggml-meta: resolve multi buffer views (llama/29266)
e0ca36d0 cuda: top-k MoE should always fire (llama/28432)
d6075ffc sycl : support new UT case for mul_mat_hadamard fp16 (llama/29218)
6570b79e sycl: extend MMVQ GLU fusion, add rms_norm+scale and ssm_conv+silu fusions (llama/28931)
a1c7f973 sycl : support op get_rows_back, only support fp32/fp16 (llama/25266)
fbcb94d0 vulkan: hide internal symbols to prevent duplicate-dlopen state destruction (llama/29139)
b18bad0d hex-dma: introduce direct-mapped DMA cache that is better suited for HVX FA mask handling (llama/29282)
e59366ba HIP : optimize IQ2/IQ3 (__vsub4 __vcmpne4) using SWAR (llama/27962)
736efcf9 opencl: add bin kernel kernel_gemm_noshuffle_q4_k_q8_1_dp4a_ila_a8_bin (llama/29056)
e900c888 vulkan: add Intel Xe flash attention optimization kernels (2/3, Xe-LPG Plus/Xe2/Xe3) (llama/24406)
ad0058be metal : gate mul_mm_id src1 rescale behind ggml_prec (llama/29029)
ff565cf3 ggml : IQ1_M build prefix sums once per block (llama/28706)
0409ed26 Performance tune for gemma4-26b-a4b flash attention shape. (llama/28450)
dd67678e opencl: add A8 Q4_0 non-MoE dp4a binary kernel (llama/29055)
898392b7 hexagon: new HMX-optimized GATED_DELTA_NET (llama/29199)
63412d34 metal : fix mask bounds in flash attention block pre-pass (llama/29220)
0e640a18 cuda: fix sm_70 tile compilation error (llama/29224)
9a7d43df ggml-cuda : convert contiguous tensors four elements at a time (llama/29155)
86ef1b15 cuda : accelerate conv2d with implicit GEMM (llama/29135)
d18a4222 ggml : fix dimension and stride truncation in ggml_permute (llama/29227)
eb279cf6 sycl : support gated DSV4_HC_PRE and optional HC_POST comb matrix (llama/29132)
0ecf57d7 sycl : pinned memory use right device context instead of 0 (llama/28895)
34a4c124 CUDA: Follow up of [#25635], refactoring FA shared smem swizzle (llama/28536)
b7b1fe47 ggml-metal : simplify fusion pattern op list declaration (llama/29206)
42c873a2 sycl : coalesce MKL-FA softmax loads instead of one work-item per row (llama/28918)
b0c51495 ggml-cpu: ARM Repack kernels for Q1_0 (llama/23492)
b50d5c31 hexagon: overhaul of buffer and DMA handling to support 64bit mappings + improvements (llama/29197)
76d02a12 cuda : tune MMVQ to MMQ crossover for SM70 (Volta) (llama/28912)
87b6876b metal : fix deprecation warnings from macOS 27 SDK (llama/29136)
29e710c4 webgpu : add fused gdn + cpy (llama/28976)
4ef9fd70 CUDA: tune FA for Gemma 4 on Ampere or newer (llama/29152)
984e400c metal : support arbitrary hc in dsv4_hc_pre (llama/29169)
3d949a36 CUDA: enable sparse fa for qwen4 (llama/28770)
c77f6bbe metal: add F16 input to the FWHT (llama/29094)
7f5ac73e hexagon: enable I32 GET_ROWS (llama/29116)
232d7189 hexagon: add support for GEGLU_QUICK (llama/29114)
099090cf hexagon: enable support for TOP_K op (llama/29113)
a00fa32e metal : add MoE and SSM_CONV fusion optimizations (llama/28948)
47456f6f metal : fix FA support checks (llama/29122)
4f1cce14 metal : support qwen4exp hc ops (llama/29000)
58844b02 cuda : fix CUB argsort corruption caused by in-place keys (llama/28389)
97dc0171 opencl: add support for bin kernel flash_attn_f32_f16_bin (llama/29046)
b29b4398 hexagon: add ROLL op support (llama/29105)
e2358dff hexagon: im2col update (llama/29103)
c01abcec hexagon: HMX flash-attention head_dim padding (support DK=DV=72) (llama/26539)
1b6c6a95 opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin (llama/28678)
8f386039 ggml-cpu: add F16 input to the FWHT (llama/27779)
bdf289e0 ggml-webgpu: fix supports_op condition for GET_ROWS (llama/28978)
26d6dcfc ggml : handle graph buffer reservation failure (llama/26070)
8e336cd0 vulkan: add IQ3_S MMQ matmul kernels (llama/28822)
6b752e5e vulkan: raise the hoisted row-id limit for mul_mat_id from 256 to 512 experts (llama/28501)
f455712a openvino : Update OpenVINO to 2026.4;fix clangd,MSVC warnings; (llama/29009)
fbdbbc78 vulkan: split buffers and debug code into separate files, add shared headers (llama/28732)
78264220 gguf : align the data section relative to the GGUF start, not the file (llama/28993)
b9e5f3ac sycl : fix the B70 mem allocate error when >19.3GB (llama/28953)
84a020e3 vulkan: skip unneeded MoE work in mul_mm coopmat1 path (llama/25483)
cb44896d sycl: ssm_conv: fuse the SiLU epilogue into the ssm_conv kernel (llama/28929)
d4b9101e opencl: fix various warnings (llama/28984)
d375e3c1 vulkan: fix buffer_reference alignment in im2col shaders (llama/28996)
31c972a9 vulkan: support qwen4exp hc ops (llama/28988)
d673fbac Fix function signature for ggml_backend_sycl_split_buffer_type (llama/28981)
67cf515a vulkan: work around NV bug with argsort_large.comp (llama/28975)
6e220ab6 hexagon: Support for K-Quants Q4_K and Q6_K (llama/28994)
6ca20bb3 hexagon: accept the zeroed rope probe in supports_op (llama/28995)
94d6e24e CUDA/HIP: improve access patterns in im2col (llama/28013)
af5153ce spacemit : fix wrong transpose function for int16 data (llama/25161)
dce52b19 rpc : invalidate cached compute graph when a referenced buffer is freed (llama/24292)
fa2c801c qwen4exp: add hc ops (llama/28901)
f9ad9868 HIP: broaden MoE ncols_opt tile heuristic on RDNA3.5 architecture (llama/28935)
0b9fb0f2 vulkan: make MUL_MAT_ID BN/2 tail unconditional (llama/28923)
edbb13e4 metal: fix NaN in mul_mm_id when activations exceed f16 range (llama/26223)
7dd0ce90 hexagon: add back missing contiguous fast-path and hvx_copy_uu for each run (llama/28886)
9e6e308e hex-cpy: use dma if src and dst are contiguous (llama/28906)
43f7501f HIP: Enable AllReduce for ROCm (llama/27825)
bae0f977 opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (llama/27637)
1fc756a2 rpc : hash-cache only weights (llama/28789)
6189e6fc cuda: support row-contiguous SUM_ROWS (llama/26308)
f69f5904 vulkan: support sparse Flash Attention (llama/28105)
4c341e28 OpenVINO: optimize stateful decode and GPU MoE inference (llama/28638)
67630d00 opencl: add generic ssm_scan (llama/28881)
4be99fe7 metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (llama/28599)
b1fd0cab cuda : enable i16 and i32 for DUP (llama/28897)
50e4eedd HIP: fattn-mma: use fp32 accumulation on MFMA devices (llama/28576)
398997ed whisper : add abort_callback on lang detection (#4077)
a44e0784 vad : reject n_encoder_layers other than 4 in model load (#4064)
307869af devops : reduce Vulkan Docker image to 39.4% of its original size (now 692 MB) (#4038)
5670d5c0 fix(yt-wsp): Resolve script path without GNU realpath (#4072)
b27fbff4 cli : load backends after validating input files (#4069)
fd7d8abb ci : update android-actions to v4.0.4 (#4074)
d5d6e59b docs : clarify VAD mode timestamps and CWD model path errors (#4019)
7a2ceef9 readme : document the ANEForge encoder backend (#4073)
4afa0094 whisper : optional ANEForge encoder backend (Apple Neural Engine) (#3905)
da545722 whisper : fix int overflow in whisper_full_parallel chunk offsets (#4044)
1d549b3c sync : ggml
ce5c557b ggml : bump version to 0.24.0 (ggml/1627)
950a4a42 tests(s390x): add non-vxe build to tests (llama/28776)
0a866737 sycl: rfc: Use radix select for top_k (llama/28670)
a3d26001 ggml-cpu : disable PCH and fix CACHE_LINE_SIZE ambiguity to fix heap corruption (llama/28882)
0bea2893 sycl : fix oneDNN scratchpad breaking the pool free order (llama/28704)
ac78ae98 ggml-cuda: fallback to F32 on device without BF16 hardware acceleration (llama/28846)
8f4cd19c ggml-cpu(s390x): guard VXE-only repack helpers (llama/28775)
51ee9279 sycl : Fix get mem error (llama/28227)
f59047cf vulkan: workaround NV queuesubmit driver bug (llama/28830)
17e63925 opencl: apply the noshuffle row-alignment rule to q4_K, q5_K and q8_0, not just q6_K (llama/28575)
8751eaeb ggml-cuda: hip add specific config table for AMD GCN (llama/27841)
24035835 syscl : Handle (fail gracefully) unsupported tq1_0 quants (llama/28681)
df31856c rpc : fix linking when compiling with BUILD_SHARED_LIBS=OFF (llama/28492)
07825dc1 opencl: fix several bugs where the backend aborts (llama/27630)
ec82a96f opencl: add bin kernel kernel_gemm_noshuffle_q4_k_f32_32b_trans_ila_a8_bin (llama/28677)
b0e076d6 webgpu: align tensor bindings to the type block size (llama/28382)
ea3ef8f5 hexagon: support for multi-device model split (aka row-split) (llama/28589)
b1779413 ggml-webgpu: Update to a recent version of Dawn (llama/28683)
2c1b5257 ggml: skip 0-sized ids tensor when offloading selected experts (llama/28739)
5f310fa1 metal : skip the empty half of the mul_mm_id token tile (llama/28301)
76da3529 cmake : add PCH and unity build to improve build times (llama/28091)
9d70ac2f metal : single-source fusion table + fusion debug rework (llama/28164)
59cca2c7 metal : fix idle threads in the remaining iq mul_mv kernels for ne00 < 1024 (llama/28692)
61b318f4 CUDA/HIP: Flash Attention tuning (gfx1201) (llama/28102)
468710cc vulkan: fix data race and OOB access in argsort(large) (llama/28705)
e5369a69 opencl: add A8 Q4_0 mm binary kernel support (llama/28268)
14e868eb vulkan: use CPU writes in ggml_backend_vk_cpy_tensor_async if the context is idle (llama/28618)
fb8f4274 vulkan: use add_alloc_dep to enable topk_moe fusion for prefill (llama/28422)
e955658b vulkan: small M matrix optimizations for qwen (llama/28457)
d374147d vulkan: fall back to shared-memory reduction for dmmv on PowerVR (llama/28341)
a6e85dd5 vulkan : add command-buffer debug labels for GPU profilers (llama/28101)
9f7331f7 ggml-cpu(s390x): add repack support for q4_0 (llama/28667)
d3945bf9 ggml-cpu(s390x): add Q1_0 vector intrinsic support (llama/28606)
3a54d53e vulkan: use spec constant for matrix matrix multiplication A-type (llama/25773)
bafeaca8 hexagon: rope updates (llama/28628)
84b8db90 vulkan: Convert FILL to distribute workgroups in 2D to avoid exceeding maxComputeWorkGroupCount (llama/28592)
f55b67f0 CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (llama/28552)
4d506f58 CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (llama/28079)
dca2df42 vulkan: add dedicated iq4_xs mat-vec shader (llama/28426)
355d90b4 vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 (llama/27471)
facf4b50 Add IQ type handling for MoE (llama/28476)
32a749dc ggml : fix msvc+clang ggml_vld1q_u32 (llama/28284)
006e53e4 Revert "ggml-cuda : restore prop.integrated on HIP builds (llama/24233)" (llama/28604)
5b984949 metal : fix idle threads in mul_mv_iq3_xxs for ne00 < 1024 (llama/28086)
8bae0820 llama : add missing headers (llama/28566)
b1275431 vulkan : fuse UNARY(GELU|SIGMOID|SILU|SOFTPLUS) + MUL (llama/27220)
69fcec3b Fix Vulkan-Hpp handle usage on 32-bit targets. (llama/22892)
37b210d8 opencl: properly handle non-contiguous inputs to conv2d (llama/28503)
33cadde1 ggml : update ggml_prec specification (llama/26675)
707c3ee1 hexagon: add RELU and LEAKY_RELU ops (llama/28585)
34f53363 Revert "CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (#24546)" (llama/28551)
0f9591a1 webgpu: format the GET_ROWS case block (llama/28542)
b61186de sycl: add a batched L2_NORM kernel (llama/28222)
8ce432f1 vulkan: add DeepSeek-V4 hyper-connection fused ops (DSV4_HC_COMB/PRE/POST) (llama/26578)
4c5a9a02 CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (llama/24546)
f15e1a67 ggml: add gfx90c HIP support (llama/26454)
a11d16a8 CUDA: branchless Q4_K/Q5_K unpack to speed up mmvq, L2 prefetch on DGX Spark (llama/26705)
60e475f2 vulkan: support type-aligned GET_ROWS (llama/28253)
c6135ac2 ggml-cuda: fix divergent barrier in f16 flash attention (llama/27870)
1da558c5 ggml: allow backend inputs to not create another split (llama/28387)
8d3ed200 vulkan: rms_norm fusion opportunities (llama/28024)
088b3b4a vulkan: add TQ1_0 support (mm, mat-vec, mat-vec-id, dequant, get_rows) (llama/27765)
ad344d16 opencl: properly choose weights pack for q4_K, q5_K mul_mat (llama/28402)
a0e75c84 cuda: fixes races in mmid and mmf (llama/28475)
1299eb6c metal : add remaining fa-vec tunings for M2 Max (llama/28458)
614c74ab metal : fix memory leak in early return (llama/28399)
0006e3e2 sycl : fix test-backend-ops CI break && restore Kronecker product FWHT support (#28016) (llama/28254)
07796107 sycl: attribute device allocations by site (GGML_SYCL_MEMTRACE) (llama/27631)
ad218df4 metal : add remaining fa-vec tunings for M3 (llama/28396)
25c6ab5d opencl: extend the elementwise and data‐movement op coverage (#27633)
7fb2dafc opencl: add Adreno xmem SDPA path (llama/26331)
70598eee scripts : fix sync (#0)
f133970b ci : use devlab-dispatch for npu-amd-windows (#4060)
1da4dc82 ci : rename cublas to cuda in release.yml (#4057)
02612981 ci : use devlab-dispatch for npu-amd-linux (#4056)