Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Official inference framework for 1-bit LLMs
AIMET is a library that provides advanced quantization and compression
Z80-μLM is a 2-bit quantized language model
Oobabooga - The definitive Web UI for local AI, with powerful features
Accessible large language models via k-bit quantization for PyTorch
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
NeurIPS2025 Spotlight] Quantized Attention
Running a 28.9M parameter LLM on an $8 microcontroller
ChatGLM3 series: Open Bilingual Chat LLMs | Open Source Bilingual Chat
Open Source Document Management System for Digital Archives
Low-code framework for building custom LLMs, neural networks
100–200× Acceleration for Video Diffusion Models
Official implementation of Watermark Anything with Localized Messages
High-performance Inference and Deployment Toolkit for LLMs and VLMs
A library for accelerating Transformer models on NVIDIA GPUs
BitNet: Scaling 1-bit Transformers for Large Language Models
Capable of understanding text, audio, vision, video
A state-of-the-art open visual language model
TechNews365 OS Admin AI intègre un Assistant Vocal IA 100% local !
A graphical manager for ollama that can manage your LLMs
PyTorch library of curated Transformer models and their components
An easy-to-use LLMs quantization package with user-friendly apis
Open platform for training, serving, and evaluating language models
Visual Instruction Tuning: Large Language-and-Vision Assistant