Showing 71 open source projects for "moe"

View related business solutions
  • Paessler: Easy to Use With Enterprise Power. Free Trial Icon
    Paessler: Easy to Use With Enterprise Power. Free Trial

    A low-code dashboard makes monitoring intuitive for any admin, while scripting and custom sensors give experts full control.

    You shouldn't have to choose between a monitoring tool that's easy to use and one that's powerful enough for a complex environment. PRTG's low-code interface lets any admin build dashboards, set alerts and monitor devices without scripting, while custom sensors and full API access are there when your team needs deeper control. One platform, no compromise. Download a free 30-day trial now.
    Get Free Download
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • 1
    MiMo-V2.5-Pro

    MiMo-V2.5-Pro

    Flagship MoE model for long-context agents and complex coding

    MiMo-V2.5-Pro is Xiaomi’s flagship Mixture-of-Experts (MoE) model built for the most demanding agentic, software engineering, and long-horizon reasoning tasks. It features approximately 1.02 trillion total parameters with 42B activated per inference, balancing extreme capability with efficient execution. The model supports a 1 million token context window, enabling it to maintain coherence across long workflows involving thousands of tool calls and multi-step reasoning chains. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Hunyuan-A13B-Instruct

    Hunyuan-A13B-Instruct

    Efficient 13B MoE language model with long context and reasoning modes

    Hunyuan-A13B-Instruct is a powerful instruction-tuned large language model developed by Tencent using a fine-grained Mixture-of-Experts (MoE) architecture. While the total model includes 80 billion parameters, only 13 billion are active per forward pass, making it highly efficient while maintaining strong performance across benchmarks. It supports up to 256K context tokens, advanced reasoning (CoT) abilities, and agent-based workflows with tool parsing. The model offers both fast and slow thinking modes, letting users trade off speed for deeper reasoning. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Mistral Large 3 675B Instruct 2512 NVFP4

    Mistral Large 3 675B Instruct 2512 NVFP4

    Quantized 675B multimodal instruct model optimized for NVFP4

    ...It retains the same instruction-tuned behavior as the FP8 model, making it ideal for production assistants, agentic workflows, scientific tasks, and long-context enterprise systems. The model integrates a 673B-parameter MoE language backbone with a 2.5B-parameter vision encoder, enabling rich multimodal analysis across text and images. Designed for efficient deployment, it runs on a single H100 or A100 node in NVFP4 while delivering performance similar to FP8 for short- and mid-context workloads.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    Mistral Large 3 675B Instruct 2512

    Mistral Large 3 675B Instruct 2512

    Frontier-scale 675B multimodal instruct MoE model for enterprise AIMis

    ...As the instruct-tuned FP8 variant, it is optimized for reliable instruction following, agentic workflows, production-grade assistants, and long-context enterprise tasks. It incorporates a massive 673B-parameter language MoE backbone and a 2.5B-parameter vision encoder, enabling rich multimodal understanding across text and images. The model supports dozens of languages and maintains strong system-prompt adherence, making it suitable for global and structured enterprise use. Designed for high performance, it runs on a single node of B200 or H200 GPUs in FP8, and can also operate in NVFP4 mode on H100 or A100 hardware. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Veeam Data Platform v13.1 Icon
    Veeam Data Platform v13.1

    Move workloads across hypervisors and clouds with no vendor lock-in. Try VDP free today.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try Now
  • 5
    gpt-oss-120b

    gpt-oss-120b

    OpenAI’s open-weight 120B model optimized for reasoning and tooling

    GPT-OSS-120B is a powerful open-weight language model by OpenAI, optimized for high-level reasoning, tool use, and agentic tasks. With 117B total parameters and 5.1B active parameters, it’s designed to fit on a single H100 GPU using native MXFP4 quantization. The model supports fine-tuning, chain-of-thought reasoning, and structured outputs, making it ideal for complex workflows. It operates in OpenAI’s Harmony response format and can be deployed via Transformers, vLLM, Ollama, LM Studio,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    LongCat-2.0

    LongCat-2.0

    Trillion-parameter MoE model for coding and million-token reasoning

    LongCat-2.0 is Meituan’s flagship open-weight Mixture-of-Experts language model designed for frontier-scale coding, reasoning, and autonomous agent workflows. It features 1.6 trillion total parameters with approximately 48 billion activated per token, combining high capability with efficient sparse inference. The model was pretrained on more than 35 trillion tokens and trained entirely on a large-scale cluster of domestically developed AI accelerators, demonstrating stable frontier-scale...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    ZAYA1-8B

    ZAYA1-8B

    Efficient MoE reasoning model for coding and math workloads

    ZAYA1-8B is a compact Mixture-of-Experts reasoning model developed by Zyphra, designed to deliver unusually high intelligence density with fewer than 1 billion active parameters. The model contains 8.4B total parameters with around 760M active during inference, allowing it to achieve strong reasoning, mathematics, and coding performance while remaining lightweight enough for efficient local or on-device deployment. ZAYA1-8B is optimized for long-form reasoning and test-time compute...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    GLM-5.3-Flash

    GLM-5.3-Flash

    Efficient 320B multimodal MoE model for coding and autonomous agents

    GLM-5.3-Flash is Z.ai’s natively multimodal model designed for efficient coding, agentic engineering, reasoning, and long-context workloads. It uses a sparse architecture with 320B total parameters and only 18B active parameters, targeting high capability with substantially lower inference costs. The model introduces a hybrid architecture combining sparse and linear attention to reduce long-context serving costs while retaining precise understanding across large inputs. It also uses...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Qwen3.8-2.4T-A95B

    Qwen3.8-2.4T-A95B

    Massive 2.4T MoE model for coding, agents, research, and reasoning

    Qwen3.8-2.4T-A95B is Qwen’s largest open-weight model and the first Qwen-Max-class model released openly, targeting advanced coding, professional work, research, and long-horizon agentic tasks. It uses a massive Mixture-of-Experts architecture with 2.4 trillion total parameters while activating 95B per token, combining Gated DeltaNet and attention layers across 512 experts. The model emphasizes reliable autonomous execution, including stronger planning, environment feedback handling, and...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Save Up to 91% on Cloud Compute With Spot VMs Icon
    Save Up to 91% on Cloud Compute With Spot VMs

    Automatic sustained-use discounts. One free VM per month. No negotiation needed.

    Run batch jobs at 60-91% off with Spot VMs. Long-running workloads get automatic discounts with sustained use.
    Start Free
  • 10
    Inkling-Small

    Inkling-Small

    Efficient multimodal MoE model for coding, tools, and reasoning

    Inkling-Small is an open-weight general-purpose multimodal model from Thinking Machines Lab, designed for agentic systems, coding assistants, chatbots, retrieval workflows, and natural-language applications. It accepts text, images, and audio as input and produces text output, with multilingual and multi-programming-language capabilities. The model uses a sparse Mixture-of-Experts architecture with 276B total parameters and 12B active per token, enabling strong performance with lower...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Inkling

    Inkling

    Frontier multimodal MoE model for coding and AI agent workflows

    Inkling is Thinking Machines Lab’s first open-weight flagship multimodal Mixture-of-Experts model, designed for advanced reasoning, coding, and autonomous agent workflows. It contains 975B total parameters with 41B active parameters per token, balancing frontier-level capability with efficient sparse inference. The model natively processes text, images, audio, and video within a unified architecture and supports an exceptionally large 1 million token context window for long-document...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Command A+

    Command A+

    4-bit Command A+ model for enterprise agents and multilingual tasks

    ...It supports text and image inputs, generates text outputs, and uses a sparse Mixture-of-Experts Transformer architecture with 218B total parameters and 25B active parameters. The W4A4 release applies 4-bit weight and activation quantization mainly to MoE experts, preserving attention components at full precision to reduce quality loss while improving speed, latency, and hardware efficiency. Cohere recommends W4A4 for most users because it offers a smaller hardware footprint with negligible benchmark differences compared to BF16 and FP8 versions. The model supports a 128K input context and 64K output length, covers 48 languages, and includes conversational tool-use capabilities with JSON-schema tools and optional citation grounding.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    DeepSeek-V4-Pro

    DeepSeek-V4-Pro

    Flagship MoE model for advanced reasoning, coding, and agents

    DeepSeek-V4-Pro is a flagship open-weight Mixture-of-Experts language model designed for high-performance reasoning, coding, and agent-based workflows at scale. It features approximately 1.6 trillion total parameters with around 49B activated during inference, enabling strong efficiency while maintaining frontier-level capability. The model supports an ultra-long context window of up to 1 million tokens, making it highly suitable for long-document reasoning, large codebases, and complex...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    DeepSeek-V4-Flash

    DeepSeek-V4-Flash

    Efficient MoE model for million-token reasoning and coding

    DeepSeek-V4-Flash is a preview Mixture-of-Experts language model built for efficient million-token context intelligence. It has 284B total parameters with 13B activated and supports a 1M-token context window, making it suitable for long-document reasoning, complex coding, agentic workflows, and large-scale information processing. The model uses a hybrid attention architecture that combines Compressed Sparse Attention and Heavily Compressed Attention to improve long-context efficiency, while...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    GigaChat 3 Ultra

    GigaChat 3 Ultra

    High-performance MoE model with MLA, MTP, and multilingual reasoning

    GigaChat 3 Ultra is a flagship instruct-model built on a custom Mixture-of-Experts architecture with 702B total and 36B active parameters. It leverages Multi-head Latent Attention to compress the KV cache into latent vectors, dramatically reducing memory demand and improving inference speed at scale. The model also employs Multi-Token Prediction, enabling multi-step token generation in a single pass for up to 40% faster output through speculative and parallel decoding techniques. Its...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Kolibri-1

    Kolibri-1

    Efficient bilingual MoE model for reasoning, coding, RAG, and agents

    Kolibri-1 is Aleph Alpha’s open-weight bilingual Mixture-of-Experts reasoning model, designed for efficient German- and English-language AI systems. It contains 78B total parameters while activating only 3.46B per token, providing substantial model capacity with relatively low inference compute. Its 50-layer architecture uses 384 experts per layer, with six routed experts and one shared expert active during processing, alongside a 4:1 combination of sliding-window and grouped-query...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    MiMo-V2.6-Flash

    MiMo-V2.6-Flash

    Efficient 309B omnimodal MoE for coding, agents, vision, and audio

    MiMo-V2.6-Flash is Xiaomi MiMo’s efficiency-balanced open-weight omnimodal model, designed to scale reinforcement learning across coding, general agents, visual tasks, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 309B total parameters while activating only 15B per token, using 256 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    MiMo-V2.6-Pro

    MiMo-V2.6-Pro

    1T omnimodal MoE model for coding, agents, and long-horizon reasoning

    MiMo-V2.6-Pro is Xiaomi MiMo’s flagship open-weight omnimodal model, built to scale reinforcement learning toward self-improvement across coding, agents, vision, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 1.02T total parameters with 42B activated per token, using 384 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Smaug Flash

    Smaug Flash

    304B MoE model optimized for agentic coding and long-running workflows

    Smaug Flash is an open-weight agentic coding model from Abacus.AI, fine-tuned from DeepSeek-V4-Flash-0731 to improve autonomous software engineering, tool use, and long-running agent workflows. It uses a 304B-parameter Mixture-of-Experts architecture with 43 layers, 256 routed experts, six selected experts per token, and one shared expert. Its attention system combines Multi-Head Latent Attention with a sparse token indexer, while DSpark multi-token prediction provides speculative decoding...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Kimi K2.7 Code

    Kimi K2.7 Code

    Coding-focused Kimi model for long-horizon agent workflows

    Kimi K2.7 Code is a coding-focused agentic model built on Kimi K2.6, designed for long-horizon software engineering, autonomous coding workflows, and complex tool-based execution. It improves end-to-end task completion across real-world programming scenarios while reducing thinking-token usage by about 30% compared with K2.6. Architecturally, it uses a 1T-parameter Mixture-of-Experts design with 32B activated parameters, 61 layers, 384 experts, a 256K-token context window, and a MoonViT...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Mistral Large 3 675B Base 2512

    Mistral Large 3 675B Base 2512

    Frontier-scale 675B multimodal base model for custom AI training

    ...The model is engineered for reliability, long-context comprehension, and stable performance across many enterprise, scientific, and knowledge-intensive workloads. Its architecture includes a powerful language MoE and a 2.5B-parameter vision encoder, enabling multimodal understanding out of the box. Mistral Large 3 Base supports deployment on-premises using FP8 or NVFP4 formats, enabling high-performance workflows on B200, H200, H100, or A100 hardware.
    Downloads: 0 This Week
    Last Update:
    See Project