Nebius Token Factory
Nebius Token Factory is a scalable AI inference platform designed to run open-source and custom AI models in production without manual infrastructure management. It offers enterprise-ready inference endpoints with predictable performance, autoscaling throughput, and sub-second latency — even at very high request volumes. It delivers 99.9% uptime availability and supports unlimited or tailored traffic profiles based on workload needs, simplifying the transition from experimentation to global deployment. Nebius Token Factory supports a broad set of open source models such as Llama, Qwen, DeepSeek, GPT-OSS, Flux, and many others, and lets teams host and fine-tune models through an API or dashboard. Users can upload LoRA adapters or full fine-tuned variants directly, with the same enterprise performance guarantees applied to custom models.
Learn more
QwenCloud
QwenCloud is an AI-native cloud platform that provides models, tools, apps, APIs, and infrastructure for building AI agents and applications. The platform offers access to featured models across LLMs, image generation, video generation, audio, speech, and multimodal workflows. QwenCloud includes flagship models such as Qwen3.8-Max, along with models for text-to-video, image generation, text-to-speech, document insight, and agentic coding. Developers can use Try AI, API keys, production-ready docs, and integrations to experiment with models and ship applications quickly. The platform also supports enterprise needs with isolated VPCs, dedicated infrastructure, compliance certifications, stable performance, and deployment monitoring. Built for developers, AI teams, enterprises, and agent builders, QwenCloud helps organizations create, test, deploy, and scale AI-native products.
Learn more
Qwen2
Qwen2 is the large language model series developed by Qwen team, Alibaba Cloud.
Qwen2 is a series of large language models developed by the Qwen team at Alibaba Cloud. It includes both base language models and instruction-tuned models, ranging from 0.5 billion to 72 billion parameters, and features both dense models and a Mixture-of-Experts model. The Qwen2 series is designed to surpass most previous open-weight models, including its predecessor Qwen1.5, and to compete with proprietary models across a broad spectrum of benchmarks in language understanding, generation, multilingual capabilities, coding, mathematics, and reasoning.
Learn more
Qwen3.8-27B
Qwen3.8-27B is an announced 27-billion-parameter model in Alibaba’s Qwen3.8 family, positioned as the compact open-weight counterpart to the much larger Qwen3.8-Max. Qwen introduced the broader Qwen3.8 generation as a new frontier model family focused on coding, agentic work, multimodal understanding, and long-running autonomous tasks. The 27B release is intended to bring that generation to a size that is far more practical for local deployment, experimentation, fine-tuning, and integration into developer workflows. Qwen has confirmed that Qwen3.8-27B will be released with open weights, extending the company’s line of downloadable mid-sized models for users who want direct control over inference and deployment. At the time of announcement, Qwen had not yet published the model card, benchmark table, architecture details, context length, quantization options, or complete deployment guidance for the 27B checkpoint.
Learn more