ReFT: Representation Finetuning for Language Models
C++ and Python support for the CUDA Quantum programming model
Windowless WebGL for node.js
ProtoMotions is a GPU-accelerated simulation and learning framework
Extract and convert data from any document, images, pdfs, word doc
Standardized Serverless ML Inference Platform on Kubernetes
Making large AI models cheaper, faster and more accessible
ETH course - Solving PDEs in parallel on GPUs
Official mirror of libplacebo
Easily compute clip embeddings and build a clip retrieval system
Embed web technologies in applications
Faster Whisper transcription with CTranslate2
OpenVINO™ Toolkit repository
Large Language Model Text Generation Inference
Public/backup repository of the GROMACS molecular simulation toolkit
Suite of reference architectures for building GPU-accelerated vision
Unified Model Serving Framework
The free, Open Source alternative to OpenAI, Claude and others
Running a big model on a small laptop
Voice Recognition to Text Tool
High-performance CPU, GPU, and memory profiler for Python
FFmpeg implements video cropping, watermarking, transcoding
Sharp Monocular Metric Depth in Less Than a Second
High-performance Kimi Delta Attention kernels
Please do not feed the models