Diversity-driven optimization and large-model reasoning ability
Tongyi Deep Research, the Leading Open-source Deep Research Agent
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
Multi-modal large language model designed for audio understanding
State-of-the-art (SoTA) text-to-video pre-trained model
Open-source framework for intelligent speech interaction
Large Multimodal Models for Video Understanding and Editing
OCR expert VLM powered by Hunyuan's native multimodal architecture
RGBD video generation model conditioned on camera input
Miso TTS is an 8 billion, highly emotive text-to-speech model
Open-weight, large-scale hybrid-attention reasoning model
Large-language-model & vision-language-model based on Linear Attention
Qwen3-omni is a natively end-to-end, omni-modal LLM
LLM-based Reinforcement Learning audio edit model
Capable of understanding text, audio, vision, video
Chinese and English multimodal conversational language model
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
Tooling for the Common Objects In 3D dataset
CodeGeeX: An Open Multilingual Code Generation Model (KDD 2023)
Memory-efficient and performant finetuning of Mistral's models
CodeGeeX2: A More Powerful Multilingual Code Generation Model
High-Resolution Image Synthesis with Latent Diffusion Models
ChatGLM-6B: An Open Bilingual Dialogue Language Model
Easy Docker setup for Stable Diffusion with user-friendly UI