tiktoken is a fast BPE tokeniser for use with OpenAI's models
Diffusion Transformer with Fine-Grained Chinese Understanding
Fast stable diffusion on CPU and AI PC
A Systematic Framework for Interactive World Modeling
Create videos with Stable Diffusion
Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD
Automatically translates the text of a video based on a subtitle file
Retrieval and Retrieval-augmented LLMs
Visual Causal Flow
Edit videos with Claude Code
End-to-end speech processing toolkit
FAIR Sequence Modeling Toolkit 2
AutoGluon: AutoML for Image, Text, and Tabular Data
Stable Diffusion WebUI optimized for AMD GPUs with editing tools
Repo of Qwen2-Audio chat & pretrained large audio language model
LLM abstractions that aren't obstructions
Unified Multimodal Understanding and Generation Models
Open source personal AI Assistant for Linux, Windows and Mac
Open source AI VTuber platform with voice chat and Live2D avatars
HunyuanVideo: A Systematic Framework For Large Video Generation Model
Qwen3 is the large language model series developed by Qwen team
Flexible Photo Recrafting While Preserving Your Identity
Bailing is a voice dialogue robot similar to GPT-4o
Build Vision Agents quickly with any model or video provider
Chinese and English multimodal conversational language model