Showing 71 open source projects for "moe"

View related business solutions
  • Veeam Data Platform v13.1 Icon
    Veeam Data Platform v13.1

    Move workloads across hypervisors and clouds with no vendor lock-in. Try VDP free today.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try Now
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    Qwen3

    Qwen3

    Qwen3 is the large language model series developed by Qwen team

    Qwen3 is a cutting-edge large language model (LLM) series developed by the Qwen team at Alibaba Cloud. The latest updated version, Qwen3-235B-A22B-Instruct-2507, features significant improvements in instruction-following, reasoning, knowledge coverage, and long-context understanding up to 256K tokens. It delivers higher quality and more helpful text generation across multiple languages and domains, including mathematics, coding, science, and tool usage. Various quantized versions,...
    Downloads: 12 This Week
    Last Update:
    See Project
  • 2
    MiMo-V2-Flash

    MiMo-V2-Flash

    MiMo-V2-Flash: Efficient Reasoning, Coding, and Agentic Foundation

    MiMo-V2-Flash is a large Mixture-of-Experts language model designed to deliver strong reasoning, coding, and agentic-task performance while keeping inference fast and cost-efficient. It uses an MoE setup where a very large total parameter count is available, but only a smaller subset is activated per token, which helps balance capability with runtime efficiency. The project positions the model for workflows that require tool use, multi-step planning, and higher throughput, rather than only single-turn chat. Architecturally, it highlights attention and prediction choices aimed at accelerating generation while preserving instruction-following quality in complex prompts. ...
    Downloads: 5 This Week
    Last Update:
    See Project
  • 3
    Kimi K2.5

    Kimi K2.5

    Moonshot's most powerful AI model

    Kimi K2.5 is Moonshot AI’s open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed vision and text tokens. Based on a 1T-parameter Mixture-of-Experts (MoE) architecture with 32B activated parameters, it integrates advanced language reasoning with strong visual understanding. K2.5 supports both “Thinking” and “Instant” modes, enabling either deep step-by-step reasoning or low-latency responses depending on the task. Designed for agentic workflows, it features an Agent Swarm mechanism that decomposes complex problems into coordinated sub-agents executing in parallel. ...
    Downloads: 17 This Week
    Last Update:
    See Project
  • 4
    Qwen3-VL

    Qwen3-VL

    Qwen3-VL, the multimodal large language model series by Alibaba Cloud

    ...It represents a major upgrade in the Qwen lineup, with stronger text generation, deeper visual reasoning, and expanded multimodal comprehension. The model supports dense and Mixture-of-Experts (MoE) architectures, making it scalable from edge devices to cloud deployments, and is available in both instruction-tuned and reasoning-enhanced variants. Qwen3-VL is built for complex tasks such as GUI automation, multimodal coding (converting images or videos into HTML, CSS, JS, or Draw.io diagrams), long-context reasoning with support up to 1M tokens, and comprehensive video understanding. ...
    Downloads: 5 This Week
    Last Update:
    See Project
  • PRTG Catches Network Issues Before They Cause Downtime Icon
    PRTG Catches Network Issues Before They Cause Downtime

    Threshold-based alerts flag problems early, so your team can act before users notice, not after.

    Reactive troubleshooting usually means hearing about a problem from frustrated users, not your monitoring tool. PRTG sets threshold-based alerts across devices, servers and applications, notifying your team by email, SMS or push the moment a metric crosses a set limit. That means catching a failing disk or overloaded server before it becomes an outage and getting time back from firefighting. Start a free trial and set your first alerts today.
    Download 30-Day Trial
  • 5
    Qwen3-Coder

    Qwen3-Coder

    Qwen3-Coder is the code version of Qwen3

    Qwen3-Coder is the latest and most powerful agentic code model developed by the Qwen team at Alibaba Cloud. Its flagship version, Qwen3-Coder-480B-A35B-Instruct, features a massive 480 billion-parameter Mixture-of-Experts architecture with 35 billion active parameters, delivering top-tier performance on coding and agentic tasks. This model sets new state-of-the-art benchmarks among open models for agentic coding, browser-use, and tool-use, matching performance comparable to leading models...
    Downloads: 14 This Week
    Last Update:
    See Project
  • 6
    Xtuner

    Xtuner

    A Next-Generation Training Engine Built for Ultra-Large MoE Models

    Xtuner is a large-scale training engine designed for efficient training and fine-tuning of modern large language models, particularly mixture-of-experts architectures. The framework focuses on enabling scalable training for extremely large models while maintaining efficiency across distributed computing environments. Unlike traditional 3D parallel training strategies, XTuner introduces optimized parallelism techniques that simplify scaling and reduce system complexity when training massive...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 7
    DualPipe

    DualPipe

    A bidirectional pipeline parallelism algorithm

    DualPipe is a bidirectional pipeline parallelism algorithm open-sourced by DeepSeek, introduced in their DeepSeek-V3 technical framework. The main goal of DualPipe is to maximize overlap between computation and communication phases during distributed training, thus reducing idle GPU time (i.e. “pipeline bubbles”) and improving cluster efficiency. Traditional pipeline parallelism methods (e.g. 1F1B or staggered pipelining) leave gaps because forward and backward phases can’t fully overlap...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    GLM-4.5

    GLM-4.5

    GLM-4.5: Open-source LLM for intelligent agents by Z.ai

    GLM-4.5 is a cutting-edge open-source large language model designed by Z.ai for intelligent agent applications. The flagship GLM-4.5 model has 355 billion total parameters with 32 billion active parameters, while the compact GLM-4.5-Air version offers 106 billion total parameters and 12 billion active parameters. Both models unify reasoning, coding, and intelligent agent capabilities, providing two modes: a thinking mode for complex reasoning and tool usage, and a non-thinking mode for...
    Downloads: 24 This Week
    Last Update:
    See Project
  • 9
    gpt-oss

    gpt-oss

    gpt-oss-120b and gpt-oss-20b are two open-weight language models

    gpt-oss is OpenAI’s open-weight family of large language models designed for powerful reasoning, agentic workflows, and versatile developer use cases. The series includes two main models: gpt-oss-120b, a 117-billion parameter model optimized for general-purpose, high-reasoning tasks that can run on a single H100 GPU, and gpt-oss-20b, a lighter 21-billion parameter model ideal for low-latency or specialized applications on smaller hardware. Both models use a native MXFP4 quantization for...
    Downloads: 4 This Week
    Last Update:
    See Project
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • 10
    DeepSpeed

    DeepSpeed

    Deep learning optimization library: makes distributed training easy

    ...Achieve extreme compression for an unparalleled inference latency and model size reduction with low costs DeepSpeed offers a confluence of system innovations, that has made large scale DL training effective, and efficient, greatly improved ease of use, and redefined the DL training landscape in terms of scale that is possible. These innovations such as ZeRO, 3D-Parallelism, DeepSpeed-MoE, ZeRO-Infinity, etc. fall under the training pillar.
    Downloads: 7 This Week
    Last Update:
    See Project
  • 11
    Grok-1

    Grok-1

    Open-source, high-performance Mixture-of-Experts large language model

    ...Due to its substantial size, utilizing Grok-1 requires a machine with significant GPU memory. The repository's MoE layer implementation prioritizes correctness over efficiency, avoiding the need for custom kernels. This is a full repo snapshot ZIP file of the Grok-1 code.
    Downloads: 19 This Week
    Last Update:
    See Project
  • 12
    yura.net
    A place for small java projects brought to you by yura dot net, currently includes: - Grasshopper - a java crash report tool for desktop, MOE iOS and android - FreeformButtonPanel - a swing component for flexible shape buttons - YIV - yura dot net Image Viewer written in java
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    DeepSeek MoE

    DeepSeek MoE

    Towards Ultimate Expert Specialization in Mixture-of-Experts Language

    DeepSeek-MoE (“DeepSeek MoE”) is the DeepSeek open implementation of a Mixture-of-Experts (MoE) model architecture meant to increase parameter efficiency by activating only a subset of “expert” submodules per input. The repository introduces fine-grained expert segmentation and shared expert isolation to improve specialization while controlling compute cost.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 14
    LLaMA-MoE

    LLaMA-MoE

    Building Mixture-of-Experts from LLaMA with Continual Pre-training

    LLaMA-MoE is an open-source project that builds mixture-of-experts language models from LLaMA through expert partitioning and continual pre-training. The repository is centered on making MoE research more accessible by offering smaller and more affordable models with only about 3.0 to 3.5 billion activated parameters, which helps reduce deployment and experimentation costs.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Markup Object Events (MOE) provides XML developers a set of Java interfaces and utility classes for working with markup both as events and as trees. MOE integrates tightly with SAX processing, while providing a basic annotated tree structure.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Moe Music is a web front end to mpg123. it uses a mySQL database as a backend to store locationsof mp3's and information. It is intended for use at LAN parties and the like to provide a central source of tunes. It works quite well so far, but it needs wor
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    Nemotron 3

    Nemotron 3

    Large language model developed and released by NVIDIA

    ...It is the post-trained and FP8-quantized variant of the Nemotron 3 Nano model, meaning its weights and activations are represented in 8-bit floating point (FP8) to dramatically reduce memory usage and computational cost while retaining high accuracy. The base Nano architecture uses a hybrid Mamba-Transformer Mixture-of-Experts (MoE) design, allowing the model to activate only a small fraction of its 31.6 billion parameters per token, which improves speed and efficiency without sacrificing quality on complex queries. This configuration supports a massive context length of up to 1 million tokens, making it suitable for long-context reasoning, agentic tasks, extended dialogues, and applications like code generation or document summarization.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Ornith-1.0

    Ornith-1.0

    Open reasoning model for agentic coding and tool workflows

    Ornith-1.0 is a large open-source reasoning model from DeepReinforce, built for agentic coding, tool use, and complex software engineering workflows. It is part of the Ornith 1.0 family, which includes dense and MoE models post-trained on Gemma 4 and Qwen 3.5. The model focuses on coding-agent performance across benchmarks such as Terminal-Bench, SWE-Bench, NL2Repo, OpenClaw, and ClawEval. Its training uses a self-improving reinforcement learning framework that optimizes not only solution attempts but also the scaffolds that guide those attempts, helping the model discover better search paths and produce higher-quality solutions. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Hy4 preview

    Hy4 preview

    770B MoE model for coding, research, reasoning, and long-context work

    Hy4 preview is Tencent’s open-weight flagship Mixture-of-Experts language model designed for advanced reasoning, software engineering, productivity, scientific research, and long-horizon tasks. It contains 770B backbone parameters while activating 49B per token across 78 layers, with 256 routed experts and one shared expert in each MoE layer. Its architecture uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse-index reuse and identity Hyper-Connections to improve information flow. A native 10B-parameter Multi-Token Prediction layer enables speculative decoding for faster inference. Hy4-preview supports a native 1M-token context window, allowing it to process large codebases, numerous files, and complex extended workflows. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Ling 3.0 Tiny

    Ling 3.0 Tiny

    Lightweight MoE model for local reasoning, coding, and AI agents

    ...It contains 7.9B total parameters while activating only 1.3B per token, using a hybrid architecture that alternates Kimi Delta Attention and Multi-Head Latent Attention with a sparse 128-expert MoE. The model supports both fast responses and configurable multi-step thinking, covering general agents, coding, mathematics, scientific reasoning, and instruction following. It is specifically optimized for local and resource-constrained deployment and has been validated on NVIDIA DGX Spark, Apple Silicon MacBooks, and Mac mini systems. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Nemotron 3.5 Lightning

    Nemotron 3.5 Lightning

    Efficient 30B MoE model for long-running agents and local inference

    NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4 is an open large language model optimized for efficient autonomous agents, sub-agent deployments, and local inference. It uses a hybrid Mixture-of-Experts architecture combining Mamba-2, MoE, and selected attention layers, with 30B total parameters but only 3B active during inference. The model supports context windows up to 1 million tokens, enabling long-running workflows and large-context reasoning. Its NVFP4 quantization reduces deployment requirements while targeting NVIDIA hardware ranging from DGX Spark and RTX 5090 systems to H100, H200, and GB200 accelerators. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Hy3 preview

    Hy3 preview

    Efficient MoE model for reasoning, coding, and AI agent workflows

    ...The model features 295B total parameters with only 21B activated during inference, plus a dedicated 3.8B Multi-Token Prediction (MTP) layer that accelerates generation through speculative decoding. Architecturally, it uses 192 routed experts with top-8 activation, a dense-MoE hybrid design, and a native 256K-token context window. Hy3-preview is optimized for efficient deployment while maintaining strong benchmark performance across reasoning, coding, and agent evaluations. It supports function calling, integration with popular agent frameworks such as OpenClaw and OpenCode, and deployment through Transformers.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    gpt-oss-20b

    gpt-oss-20b

    OpenAI’s compact 20B open model for fast, agentic, and local use

    GPT-OSS-20B is OpenAI’s smaller, open-weight language model optimized for low-latency, agentic tasks, and local deployment. With 21B total parameters and 3.6B active parameters (MoE), it fits within 16GB of memory thanks to native MXFP4 quantization. Designed for high-performance reasoning, it supports Harmony response format, function calling, web browsing, and code execution. Like its larger sibling (gpt-oss-120b), it offers adjustable reasoning depth and full chain-of-thought visibility for better interpretability. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Qwen3.8-Flash-Next

    Qwen3.8-Flash-Next

    Efficient multimodal MoE model for coding, reasoning, and AI agents

    ...Its hybrid architecture combines Gated DeltaNet with Qwen Sparse Attention (QSA), which processes micro-blocks rather than individual tokens to reduce latency in long-context agent workloads. The model also introduces Gated Residual connections and scalable n-gram embeddings to improve efficiency while limiting inference overhead. It contains 512 MoE experts, activating 10 routed experts plus one shared expert per token. Qwen3.8-Flash-Next natively handles text, images, and video and supports a 262K-token context window extensible to 1M tokens. It targets coding, tool use, professional tasks, computer interaction, multimodal reasoning, and long-horizon agents, with configurable thinking modes and reasoning effort.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    Solar Open 2

    Solar Open 2

    Efficient 250B MoE model for agents, coding, and long-context work

    Solar Open 2 is Upstage’s 250B-A15B open-weight large language model designed for agentic workflows, office productivity, document-intensive tasks, coding, and reasoning. Its Hybrid-Attention Mixture-of-Experts architecture contains 250B total parameters while activating only 15B per token, combining three linear-attention layers with one softmax-attention layer for efficient inference. The model supports a native 1M-token context window and uses NoPE instead of rotary positional encoding,...
    Downloads: 0 This Week
    Last Update:
    See Project