Download Latest Version llama-b8783-bin-ubuntu-openvino-2026.0-x64.tar.gz (78.1 MB)
Email in envelope

Get an email when there's a new version of llama.cpp

Home / b8779
Name Modified Size InfoDownloads / Week
Parent folder
llama-b8779-xcframework.zip 2026-04-13 178.9 MB
llama-b8779-bin-win-vulkan-x64.zip 2026-04-13 58.4 MB
llama-b8779-bin-win-sycl-x64.zip 2026-04-13 136.9 MB
llama-b8779-bin-win-opencl-adreno-arm64.zip 2026-04-13 33.9 MB
llama-b8779-bin-win-hip-radeon-x64.zip 2026-04-13 361.9 MB
llama-b8779-bin-win-cuda-13.1-x64.zip 2026-04-13 166.5 MB
llama-b8779-bin-win-cuda-12.4-x64.zip 2026-04-13 249.1 MB
llama-b8779-bin-win-cpu-x64.zip 2026-04-13 39.9 MB
llama-b8779-bin-win-cpu-arm64.zip 2026-04-13 32.7 MB
llama-b8779-bin-ubuntu-x64.tar.gz 2026-04-13 32.2 MB
llama-b8779-bin-ubuntu-vulkan-x64.tar.gz 2026-04-13 50.7 MB
llama-b8779-bin-ubuntu-vulkan-arm64.tar.gz 2026-04-13 42.4 MB
llama-b8779-bin-ubuntu-s390x.tar.gz 2026-04-13 35.5 MB
llama-b8779-bin-ubuntu-rocm-7.2-x64.tar.gz 2026-04-13 169.5 MB
llama-b8779-bin-ubuntu-openvino-2026.0-x64.tar.gz 2026-04-13 77.7 MB
llama-b8779-bin-ubuntu-arm64.tar.gz 2026-04-13 28.4 MB
llama-b8779-bin-macos-x64.tar.gz 2026-04-13 40.9 MB
llama-b8779-bin-macos-arm64.tar.gz 2026-04-13 40.8 MB
llama-b8779-bin-macos-arm64-kleidiai.tar.gz 2026-04-13 40.8 MB
llama-b8779-bin-910b-openEuler-x86-aclgraph.tar.gz 2026-04-13 73.2 MB
llama-b8779-bin-910b-openEuler-aarch64-aclgraph.tar.gz 2026-04-13 65.6 MB
llama-b8779-bin-310p-openEuler-x86.tar.gz 2026-04-13 73.2 MB
llama-b8779-bin-310p-openEuler-aarch64.tar.gz 2026-04-13 65.6 MB
cudart-llama-bin-win-cuda-13.1-x64.zip 2026-04-13 402.6 MB
cudart-llama-bin-win-cuda-12.4-x64.zip 2026-04-13 391.4 MB
b8779 source code.tar.gz 2026-04-13 33.8 MB
b8779 source code.zip 2026-04-13 35.0 MB
README.md 2026-04-13 3.5 kB
Totals: 28 Items   3.0 GB 3
vulkan: Flash Attention DP4A shader for quantized KV cache (#20797) * use integer dot product for quantized KV flash attention * small improvements * fix SHMEM_STAGING indexing * add missing KV type quants * fixes * add supported quants to FA tests * readd fast paths for <8bit quants * fix mmq gate and shmem checks

macOS/iOS:

Linux:

Windows:

openEuler:

Source: README.md, updated 2026-04-13