Qwen3-Omni is a natively end-to-end multilingual omni-modal foundation model that processes text, images, audio, and video and delivers real-time streaming responses in text and natural speech. It uses a Thinker-Talker architecture with a Mixture-of-Experts (MoE) design, early text-first pretraining, and mixed multimodal training to support strong performance across all modalities without sacrificing text or image quality. The model supports 119 text languages, 19 speech input languages, and 10 speech output languages. It achieves state-of-the-art results: across 36 audio and audio-visual benchmarks, it hits open-source SOTA on 32 and overall SOTA on 22, outperforming or matching strong closed-source models such as Gemini-2.5 Pro and GPT-4o. To reduce latency, especially in audio/video streaming, Talker predicts discrete speech codecs via a multi-codebook scheme and replaces heavier diffusion approaches.

Features

  • Processes and understands text, images, audio, and video as inputs in mixed or separate forms
  • Generates real-time responses both as text and natural speech (audio output)
  • Multilingual capabilities: supports 119 text languages, 19 speech input languages, 10 speech output languages
  • Comes with variants/checkpoints: e.g. Instruct (thinker + talker), Thinking (thinker only), Captioner for detailed audio captioning, etc.
  • Efficient architecture: MoE-based Thinker–Talker design, multi-codebook to reduce latency, support for FlashAttention (v2) and use of frameworks such as Transformers and vLLM
  • Deployment support: Docker image, demos (web UI), offline and online API options, detailed cookbooks for various use-cases (speech recognition, OCR, audio-visual dialogue, etc.)

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow Qwen3-Omni

Qwen3-Omni Web Site

Other Useful Business Software
MongoDB Atlas runs apps anywhere Icon
MongoDB Atlas runs apps anywhere

Deploy in 115+ regions with the modern database for every enterprise.

MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Qwen3-Omni!

Additional Project Details

Operating Systems

Linux, Mac, Windows

Programming Language

Python

Related Categories

Python Large Language Models (LLM), Python AI Models

Registered

2025-09-23