High-Fidelity and Controllable Generation of Textured 3D Assets
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Large-language-model & vision-language-model based on Linear Attention
MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training
An Efficient Agentic Model for Computer Use
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
LLM-based Reinforcement Learning audio edit model
Robust Speech Recognition Across Languages, Dialects
Tiny vision language model
Open-source industrial-grade ASR models
Project Lyra: Open Generative 3D World Models
Inference script for Oasis 500M
Generate Any 3D Scene in Seconds
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
Uncommon Objects in 3D dataset
GPT4V-level open-source multi-modal model based on Llama3-8B
Multi-modal large language model designed for audio understanding
Qwen3-omni is a natively end-to-end, omni-modal LLM
Chinese and English multimodal conversational language model
High-Resolution Image Synthesis with Latent Diffusion Models
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation
GUI shell for running local LLM on desktop
Open-source, high-performance Mixture-of-Experts large language model
Qwen2.5-Coder is the code version of Qwen2.5, the large language model
Memory-efficient and performant finetuning of Mistral's models