157 models, 30 providers, one command to find what runs on hardware
Qwen3-TTS is an open-source series of TTS models
Fast State-of-the-Art Static Embeddings
Achieving 3+ generation speedup on reasoning tasks
Ultralytics YOLO
LightLLM is a Python-based LLM (Large Language Model) inference
TokenSpeed is a speed-of-light LLM inference engine
100–200× Acceleration for Video Diffusion Models
Local long-term memory engine for AI apps with persistent storage
A Web UI for easy subtitle using whisper model
High-Quality Voice Cloning TTS for 600+ Languages
Mooncake is the serving platform for Kimi
Deep learning optimization library: makes distributed training easy
Ultra-Efficient LLMs on End Device
Multilingual speech recognition and audio understanding model
The best agent harness
MiMo-V2-Flash: Efficient Reasoning, Coding, and Agentic Foundation
This repo contains the code for 1D tokenizer and generator
The repository provides code for running inference with SAM 2
Industrial-level controllable zero-shot text-to-speech system
Sharp Monocular Metric Depth in Less Than a Second
Image generation model with single-stream diffusion transformer
An Open Source text-to-speech system built by inverting Whisper
Apple Intelligence from the command line
Running a big model on a small laptop