23 Integrations with Alibaba Cloud Model Studio

View a list of Alibaba Cloud Model Studio integrations and software that integrates with Alibaba Cloud Model Studio below. Compare the best Alibaba Cloud Model Studio integrations as well as features, ratings, user reviews, and pricing of software that integrates with Alibaba Cloud Model Studio. Here are the current Alibaba Cloud Model Studio integrations in 2026:

  • 1
    Qwen3.8-Max
    Qwen3.8-Max is Qwen’s most capable model to date, built as a Max-class AI model for coding, work, research, long-horizon tasks, and multimodal agents. It scales to 2.4 trillion parameters with 95 billion active parameters and is available through QwenCloud. The model is designed to complete complex, open-ended tasks end to end with greater reliability and minimal human involvement. Qwen3.8-Max supports autonomous coding workflows, agentic development, research reproduction, visual reasoning, document understanding, video analysis, and real-world productivity tasks. It can integrate with popular agent frameworks and coding assistants, including Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw. Built for developers, researchers, enterprises, and AI agent builders, Qwen3.8-Max helps teams automate sophisticated work across code, documents, tools, interfaces, and multimodal content.
    Starting Price: $2 per 1M (input)
  • 2
    HappyHorse 1.1
    HappyHorse-1.1-T2V is a text-to-video generation model available through QwenCloud. The model turns text prompts into video output with improved semantic understanding, cinematic shot control, and dynamic motion rendering. HappyHorse-1.1-T2V is designed to capture creative intent more accurately while producing videos with smoother motion, richer details, and stronger visual consistency. It supports natural character actions, scene atmosphere, and physical dynamics for more realistic video generation. The model can be accessed through the QwenCloud API with configurable options such as resolution, aspect ratio, and duration. Built for developers, creators, and AI product teams, HappyHorse-1.1-T2V helps generate high-quality videos from text prompts at scale.
  • 3
    OpenAI

    OpenAI

    OpenAI

    OpenAI’s mission is to ensure that artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work—benefits all of humanity. We will attempt to directly build safe and beneficial AGI, but will also consider our mission fulfilled if our work aids others to achieve this outcome. Apply our API to any language task — semantic search, summarization, sentiment analysis, content generation, translation, and more — with only a few examples or by specifying your task in English. One simple integration gives you access to our constantly-improving AI technology. Explore how you integrate with the API with these sample completions.
  • 4
    Qwen3.6-27B
    Qwen3.6-27B is a dense, open source multimodal language model in the Qwen3.6 series, designed to deliver flagship-level performance in coding, reasoning, and agent-based workflows while maintaining a relatively efficient parameter size of 27 billion. It is positioned as a high-performance general model that “punches above its weight,” achieving results competitive with or superior to significantly larger models on key benchmarks, particularly in agentic coding tasks. It supports both thinking and non-thinking modes, allowing it to dynamically balance deep reasoning with fast responses depending on the task, and integrates capabilities across text and multimodal inputs such as images and video. Built as part of the Qwen3.6 family, the model emphasizes real-world usability, stability, and developer productivity, incorporating improvements driven by community feedback and practical deployment needs.
    Starting Price: Free
  • 5
    Qwen

    Qwen

    Alibaba

    Qwen is a powerful, free AI assistant built on the advanced Qwen model series, designed to help anyone with creativity, research, problem-solving, and everyday tasks. While Qwen Chat is the main interface for most users, Qwen itself powers a broad range of intelligent capabilities including image generation, deep research, website creation, advanced reasoning, and context-aware search. Its multimodal intelligence enables Qwen to understand and process text, images, audio, and video simultaneously for richer insights. Qwen is available on web, desktop, and mobile, ensuring seamless access across all devices. For developers, the Qwen API provides OpenAI-compatible endpoints, making integration simple and allowing Qwen’s intelligence to power apps, services, and automation. Whether you're chatting through Qwen Chat or building with the Qwen API, Qwen delivers fast, flexible, and highly capable AI support.
    Starting Price: Free
  • 6
    Qwen3.8-27B
    Qwen3.8-27B is a compact open-weights model in Alibaba’s Qwen3.8 family, aimed at developers and researchers who want strong local AI performance without using the full Max-scale model. Reports from Alibaba’s Qwen3.8 launch state that Qwen3.8-27B was planned for open-weight release alongside Qwen3.8-Max, expanding access for builders working on AI applications. The model is positioned for coding, research, professional workflows, and local deployment scenarios where a 27B model can be more practical than frontier-scale systems. Qwen3.8’s broader launch emphasizes software development, document processing, data analysis, and professional “cowork” use cases. Qwen3.8-27B is especially relevant for teams that need a capable open model for experimentation, coding agents, assistant workflows, and self-hosted inference. Built for practical deployment, Qwen3.8-27B gives developers a smaller Qwen3.8 option for building AI tools, testing agents, and running advanced language model workflows.
  • 7
    Qwen3.6-Max-Preview
    Qwen3.6-Max-Preview is a next-generation frontier language model designed to push the limits of intelligence, instruction following, and real-world agent capabilities within the Qwen ecosystem. Building on the Qwen3 series, this preview release introduces stronger world knowledge, sharper instruction alignment, and significant improvements in agentic coding performance, enabling the model to better handle complex, multi-step tasks and software engineering workflows. It is engineered for advanced reasoning and execution scenarios, where the model not only generates responses but also interacts with tools, processes long contexts, and supports structured problem-solving across domains such as coding, research, and enterprise workflows. The architecture continues the Qwen focus on large-scale, high-efficiency models capable of handling extensive context windows and delivering consistent performance across multilingual and knowledge-intensive tasks.
    Starting Price: Free
  • 8
    Qwen3.7-Max
    Qwen3.7-Max is Qwen’s latest proprietary model designed for the agent era, built to be a versatile agent foundation that is equally capable of writing and debugging code, automating office workflows, and sustaining autonomous browser sessions over long horizons. It reaches frontier-level coding performance, with stronger results across software engineering, terminal tasks, GUI grounding, web browsing, and agentic tool use. Qwen3.7-Max is designed to reduce the gap between model intelligence and real agent execution by supporting planning, long-context reasoning, reliable function calling, and multi-step task completion across complex workflows. It also strengthens multimodal and document-oriented work through Qwen Studio, which supports chatbot interaction, image and video understanding, image generation, document processing, presentation generation, coding assistance, deep research, and web development.
    Starting Price: Free
  • 9
    Qwen3.5

    Qwen3.5

    Alibaba

    Qwen3.5 is a next-generation open-weight multimodal large language model designed to power native vision-language agents. The flagship release, Qwen3.5-397B-A17B, combines a hybrid linear attention architecture with sparse mixture-of-experts, activating only 17 billion parameters per forward pass out of 397 billion total to maximize efficiency. It delivers strong benchmark performance across reasoning, coding, multilingual understanding, visual reasoning, and agent-based tasks. The model expands language support from 119 to 201 languages and dialects while introducing a 1M-token context window in its hosted version, Qwen3.5-Plus. Built for multimodal tasks, it processes text, images, and video with advanced spatial reasoning and tool integration. Qwen3.5 also incorporates scalable reinforcement learning environments to improve general agent capabilities. Designed for developers and enterprises, it enables efficient, tool-augmented, multimodal AI workflows.
    Starting Price: Free
  • 10
    Happy Shrimp 1.0

    Happy Shrimp 1.0

    Alibaba Cloud

    Happy Shrimp 1.0 is an AI music generation model designed to turn vague ideas into fully produced songs from a single prompt. Users can start with an emotion, story concept, target genre, or creative direction, without needing technical music knowledge such as BPM, key signature, or instrumentation. The model supports text-to-music generation for complete vocal songs and instrumental compositions, producing melody, arrangement, lyrics, and vocals from scratch or using lyrics supplied by the user. Built with broad musical knowledge, it understands creative aesthetics across genres, regions, eras, and cultural contexts, translating descriptive prompts into structured musical outputs. It performs across styles including Chinese music, pop, R&B and soul, hip hop, rock, funk, electronic, classical, and jazz. Using world knowledge and music-domain reasoning, the model represents both the “grammar” of music.
    Starting Price: Free
  • 11
    Qwen3.5-Plus
    Qwen3.5-Plus is a high-performance native vision-language model designed for efficient text generation, deep reasoning, and multimodal understanding. Built on a hybrid architecture that combines linear attention with a sparse mixture-of-experts design, it delivers strong performance while optimizing inference efficiency. The model supports text, image, and video inputs and produces text outputs, making it suitable for complex multimodal workflows. With a massive 1 million token context window and up to 64K output tokens, Qwen3.5-Plus enables long-form reasoning and large-scale document analysis. It includes advanced capabilities such as structured outputs, function calling, web search, and tool integration via the Responses API. The model supports prefix continuation, caching, batch processing, and fine-tuning for flexible deployment. Designed for developers and enterprises, Qwen3.5-Plus provides scalable, high-throughput AI performance with OpenAI-compatible API access.
    Starting Price: $0.4 per 1M tokens
  • 12
    Qwen-Audio-3.0-TTS-Plus
    Qwen-Audio-3.0-TTS-Plus is the high-quality variant of Qwen-Audio-3.0-TTS, optimized for naturalness and timbre fidelity when output quality matters more than speed. It supports 16 languages, plus improved fidelity for several Chinese dialects. The model delivers strong multilingual intelligibility and ranks first in speaker similarity across all supported languages, helping cloned voices remain recognizable and consistent across linguistic contexts. Developers can direct delivery through ordinary natural-language instructions instead of manually tuning acoustic parameters, controlling emotion, role, scenario, pacing, projection, and tone with simple prompts. Inline tags provide fine-grained control over breaths, laughter, emotional shifts, and other non-verbal details, making the model useful for narration, games, character dialogue, and dubbing.
  • 13
    Qwen-Audio-3.0-TTS-Flash
    Qwen-Audio-3.0-TTS-Flash is the real-time variant of Qwen-Audio-3.0-TTS, tuned for interactive applications with first-packet latency at the 300 ms level. It supports 16 languages, along with improved fidelity for several Chinese dialects. Across multilingual evaluations, Flash delivers the lowest average WER/CER in the family at 3.87, showing strong intelligibility while preserving speaker identity across diverse languages. Developers can guide delivery with plain-language instructions instead of manually adjusting acoustic parameters, controlling emotion, role, scenario, pace, projection, and tone through simple prompts. Inline tags add precise non-verbal details, making the model well-suited to conversational agents, narration, games, dubbing, and other expressive speech experiences. Voice cloning is designed to work with imperfect reference audio; targeted acoustic simulation suppresses noise and reverberation while retaining the original speaker’s timbre.
  • 14
    Wan3.0

    Wan3.0

    Alibaba

    Wan3.0 is an all-in-one video generation model from Qwen Cloud that unifies multiple creative capabilities in a single system, including text-to-video, image-to-video, reference-to-video, editing, replication, and driving. It supports audio, image, text, and video inputs and produces video output, allowing creators to guide generation with several types of source material instead of relying on text prompts alone. The model can generate videos up to 30 seconds long and supports omni-modal reference, giving users more flexibility when carrying visual, motion, character, or other creative cues into a new result. Wan3.0 can also parse files, web pages, and complex images as part of the generation workflow. Its image-to-video capabilities include first-frame and first-and-last-frame generation, making it possible to define how a sequence begins or anchor both ends of a shot.
    Starting Price: $0.05 per second
  • 15
    Qwen3.8-Flash-Next
    Qwen3.8-Flash-Next is an open-weight multimodal Mixture-of-Experts model and an early preview of the architecture planned for Qwen4. It systematically upgrades attention, residual connections, embeddings, and optimization to improve capability, computational efficiency, model capacity, and training stability. Its hybrid architecture combines Gated DeltaNet, which efficiently compresses historical information, with Qwen Sparse Attention, which selects important context at the micro-block level to reduce attention and indexing costs on long sequences. Gated Residual widens the residual stream into four branches and dynamically controls information flow across layers, while N-gram Embedding adds large-scale local-pattern memory with very little extra per-token computation and can be offloaded to host memory. The model uses a 125B-parameter main network plus 51B N-gram embedding parameters, while activating only 6B parameters per token.
    Starting Price: $2 per 1M (input)
  • 16
    Step 5 Preview
    Step 5 Preview is StepFun’s flagship model for agentic work, designed for real-world tasks across software engineering and professional knowledge work, with particular strength in finance. It natively supports text, image, and video input and provides a 1M-token context window, enabling tasks that require large amounts of information, tool calls, and continuous progress toward a deliverable. The model can analyze long documents, multiple source materials, and conversation history for cross-document question answering and research organization. For programming and software engineering, it works across multiple languages and can support troubleshooting, code changes, verification, and test creation. Its multi-step agent capabilities let applications provide tools for retrieving information, processing documents, conducting deep research, and producing analytical reports. Multimodal understanding combines images, video, and text for chart analysis, screenshot question answering, etc.
    Starting Price: $0.04 per input
  • 17
    Alibaba Virtual Private Cloud
    VPC helps you build an isolated network environment based on Alibaba Cloud including customizing the IP address range, network segment, route table, and gateway. In addition, you can connect VPC and a traditional IDC through a leased line, VPN, or GRE to provide hybrid cloud services. The service system can be deployed in both local and on-cloud IDCs. Different service modules are built on Alibaba Cloud VPC to create fully isolated on-cloud environments. On-cloud and off-cloud services are interacted with each other through the Internet.
  • 18
    Qwen3.6-Plus
    Qwen3.6-Plus is an advanced AI model developed by Alibaba Cloud, designed to power real-world intelligent agents and complex workflows. It introduces significant improvements in agentic coding, enabling developers to handle everything from frontend development to large-scale codebase management. The model features a massive 1 million token context window, allowing it to process and reason over long and complex inputs. It integrates reasoning, memory, and execution capabilities to deliver highly accurate and reliable results. Qwen3.6-Plus also enhances multimodal capabilities, enabling it to understand and analyze images, videos, and documents. The platform is optimized for real-world applications, including automation, planning, and tool-based workflows. Overall, it provides a powerful foundation for building next-generation AI agents and intelligent systems.
  • 19
    Happy Horse
    Happy Horse is an AI video generation and editing platform that helps users turn creative ideas into cinematic videos. The platform supports video creation from text, reference inputs, and first-frame prompts, giving creators flexible ways to bring visual concepts to life. Users can also edit videos by modifying details and refining generated results. Happy Horse features a creative community showcase with short films, featured videos, and AI cinema projects. The platform includes credits for generation, promotional offers, and tools for experimenting with imaginative video concepts. Happy Horse helps creators, artists, filmmakers, and storytellers capture ideas quickly and transform them into expressive AI-generated video content.
  • 20
    Qwen3.7-Plus
    Qwen3.7-Plus is a multimodal agent model that unifies vision and language into a single, versatile agent foundation. Building on Qwen3.7’s agentic intelligence, it extends Qwen’s capabilities into visual understanding, visual reasoning, grounded interaction, and multimodal tool use, enabling agents to perceive, analyze, and act across text, images, documents, screens, and complex real-world contexts. It is designed for tasks that require more than static question answering, including visual search, document comprehension, chart and table analysis, screen understanding, GUI interaction, image-grounded reasoning, and agent workflows that combine perception with planning and execution. Qwen3.7-Plus strengthens the connection between language reasoning and visual evidence, allowing users to ask questions about images, interpret dense multimodal inputs, extract structured information, and generate responses that reflect both context and visual details.
  • 21
    Qwen3.8-2.4T-A95B
    Qwen3.8-2.4T-A95B is the largest open model in the Qwen3.8 family, bringing Qwen-Max-class capabilities to an open release. Built on the architectural foundation of Qwen3.5, it delivers substantial improvements across coding, professional work, research, and long-horizon agentic tasks, with a focus on carrying complex, multi-step work through to completion more reliably. The causal language model uses a mixture-of-experts architecture with 2.4 trillion total parameters and 95 billion activated parameters, including 512 experts with 10 routed and one shared expert active at a time. It supports a native context length of 262,144 tokens that can be extended to approximately 1.01 million tokens. Agent execution is strengthened through better autonomous planning and improved handling of environment feedback, while broader compatibility with popular agent harnesses and development tools simplifies integration into existing stacks.
  • 22
    Qwen3.8-Omni-Flash
    Qwen3.8-Omni-Flash is a next-generation native omnimodal model designed to strengthen agent capabilities in real-world productivity scenarios, advancing from understanding multimodal content to planning tasks, calling tools, and completing creative work. Built on the Qwen3.8-Flash-Next architecture, it accepts text, image, audio, and video inputs with a context window of up to 1 million tokens while maintaining strong text performance. Beyond coding, knowledge work, and GUI interaction, it extends agentic workflows centered on audio and video, including video editing, music video creation, film production and commentary, audiovisual summarization, and real-time conversations. The model improves long-form audio and audiovisual understanding through controllable descriptions, agentic evidence gathering, meeting understanding, and video-centered deep research. Users can specify the subject, time range, level of detail, and output format for video analysis, enabling overviews, etc.
  • 23
    Omni

    Omni

    Omni

    Get the best of both with a more evolved analytics experience. Find and share the metrics that matter with a more reliable model your entire organization can use - and reuse. Omni takes only minutes to set up and immediately begins to learn from every query, optimizing performance over time. Our innovative BI auto-builds a data model as you query, creating instantly shareable metrics anyone can use. Easy-to-use UX allows anyone, regardless of skill level, to access, explore, and share the data they need. Omni even applies one-off queries directly into the data model, expanding and building upon your shareable data. Omni, business intelligence that starts with fluid data exploration, and matures to a structured data model that promotes the accurate and efficient use of data.
  • Previous
  • You're on page 1
  • Next