Audience
Developers, AI teams, and organizations seeking to build agents that understand multimodal content, use tools, and complete complex audio, video, and productivity workflows
About Qwen3.8-Omni-Flash
Qwen3.8-Omni-Flash is a next-generation native omnimodal model designed to strengthen agent capabilities in real-world productivity scenarios, advancing from understanding multimodal content to planning tasks, calling tools, and completing creative work. Built on the Qwen3.8-Flash-Next architecture, it accepts text, image, audio, and video inputs with a context window of up to 1 million tokens while maintaining strong text performance. Beyond coding, knowledge work, and GUI interaction, it extends agentic workflows centered on audio and video, including video editing, music video creation, film production and commentary, audiovisual summarization, and real-time conversations. The model improves long-form audio and audiovisual understanding through controllable descriptions, agentic evidence gathering, meeting understanding, and video-centered deep research. Users can specify the subject, time range, level of detail, and output format for video analysis, enabling overviews, etc.