Accurate × Fast × Comprehensive
Industrial-level controllable zero-shot text-to-speech system
Recovering the Visual Space from Any Views
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Advancing Open-source World Models
GLM-4 series: Open Multilingual Multimodal Chat LMs
Open-source multi-speaker long-form text-to-speech model
A Pragmatic VLA Foundation Model
Z80-μLM is a 2-bit quantized language model
OCR expert VLM powered by Hunyuan's native multimodal architecture
Provides convenient access to the Anthropic REST API from any Python 3
DeepSeek Coder: Let the Code Write Itself
An experimental version of DeepSeek model
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
1B text generation model based on the HRM architecture
Official implementation of Watermark Anything with Localized Messages
Open-Source Financial Large Language Models
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Sharp Monocular Metric Depth in Less Than a Second
Miso TTS is an 8 billion, highly emotive text-to-speech model
Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI
Audio foundation model excelling in audio understanding
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Fast-stable-diffusion + DreamBooth
CogView4, CogView3-Plus and CogView3(ECCV 2024)