Buzz transcribes and translates audio offline
MiniMax H3 is a general-purpose, omni-modal generative system
Awesome multilingual OCR toolkits based on PaddlePaddle
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Open-source, high-performance AI model with advanced reasoning
The most powerful local music generation model
Official inference repo for FLUX.1 models
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Official Python inference and LoRA trainer package
Open-source multi-speaker long-form text-to-speech model
Lets make video diffusion practical
From Images to High-Fidelity 3D Assets
Code for running inference and finetuning with SAM 3 model
Native and Compact Structured Latents for 3D Generation
Powerful AI language model (MoE) optimized for efficiency/performance
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Industrial-level controllable zero-shot text-to-speech system
Official inference repo for FLUX.2 models
Python inference and LoRA trainer package for the LTX-2 audio–video
Fast stable diffusion on CPU and AI PC
High-Resolution Image Synthesis with Latent Diffusion Models
Python bindings for llama.cpp
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Reference PyTorch implementation and models for DINOv3