Designed for text embedding and ranking tasks
HY-Motion model for 3D character animation generation
Large-language-model & vision-language-model based on Linear Attention
MOSS‑TTS Family open‑source speech and sound generation model
General-purpose image editing model that delivers high-fidelity
Diffusion Transformer with Fine-Grained Chinese Understanding
Visual Causal Flow
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
LLM-based Reinforcement Learning audio edit model
Phi-3.5 for Mac: Locally-run Vision and Language Models
Bidirectional token-classification model for identifiable info
Ultra-Efficient LLMs on End Device
Unified Multimodal Understanding and Generation Models
Implementation of "MobileCLIP" CVPR 2024
Qwen3 is the large language model series developed by Qwen team
Open-source framework for intelligent speech interaction
OCR expert VLM powered by Hunyuan's native multimodal architecture
CogView4, CogView3-Plus and CogView3(ECCV 2024)
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Renderer for the harmony response format to be used with gpt-oss
A Systematic Framework for Interactive World Modeling
Multimodal Diffusion with Representation Alignment
Repo of Qwen2-Audio chat & pretrained large audio language model
Audio Language Models are Few-Shot Learners