Open image model at the forefront of design
Wan2.2: Open and Advanced Large-Scale Video Generative Model
AI cognitive-enhancement Skills based on Anthropic's J-space
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
Qwen-Image is a powerful image generation foundation model
Controllable & emotion-expressive zero-shot TTS
Open-source framework for intelligent speech interaction
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Renderer for the harmony response format to be used with gpt-oss
Bidirectional token-classification model for identifiable info
Achieving 3+ generation speedup on reasoning tasks
Foundation model for image generation
A Pragmatic VLA Foundation Model
LLM-based Reinforcement Learning audio edit model
MOSS‑TTS Family open‑source speech and sound generation model
Advancing Open-source World Models
High-Fidelity and Controllable Generation of Textured 3D Assets
Real-time behaviour synthesis with MuJoCo, using Predictive Control
Let us control diffusion models
Code for "Image Generation from Scene Graphs", Johnson et al, CVPR 201
Dia-1.6B generates lifelike English dialogue and vocal expressions