A 2.78-trillion-parameter Kimi K3 running inference on a single CPU
Port of Facebook's LLaMA model in C/C++
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
Awesome multilingual OCR toolkits based on PaddlePaddle
Flux 2 image generation model pure C inference
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
C#/.NET binding of llama.cpp, including LLaMa/GPT model inference
AlphaFold 3 inference pipeline
Run the full 2.78-trillion-parameter Kimi K3 model
FAIR Sequence Modeling Toolkit 2
Open-source large language model family from Tencent Hunyuan
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
Hackable and optimized Transformers building blocks
VMZ: Model Zoo for Video Modeling
Fast, Sharp & Reliable Agentic Intelligence
Foundational Models for State-of-the-Art Speech and Text Translation
Clean and efficient FP8 GEMM kernels with fine-grained scaling
FlashMLA: Efficient Multi-head Latent Attention Kernels
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)
Runtime extension of Proximus enabling Deployment on AMD Ryzen™ AI
Distribution TN 365 KDE moderne et stable !
Real-time behaviour synthesis with MuJoCo, using Predictive Control
A fast, local neural text to speech system
llama.go is like llama.cpp in pure Golang