Showing 215 open source projects for "mixture"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • 1
    MoBA

    MoBA

    MoBA: Mixture of Block Attention for Long-Context LLMs

    MoBA, short for Mixture of Block Attention, is an open-source research implementation of a novel attention mechanism designed to improve the efficiency of large language models processing extremely long contexts. The architecture adapts ideas from Mixture-of-Experts networks and applies them directly to the attention mechanism of transformer models. Instead of forcing each token to attend to every other token in the sequence, MoBA divides the context into blocks and dynamically routes queries to only the most relevant segments of information. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Ornith-1.0

    Ornith-1.0

    Ornith-1.0 is a self-improving open-source models for agentic coding

    ...It is designed for coding agents that need to solve software engineering tasks through iterative tool use and solution rollouts. The project presents 9B dense, 31B dense, 35B mixture-of-experts, and 397B mixture-of-experts variants. These models are post-trained on top of Gemma 4 and Qwen 3.5 foundations. Its training approach uses reinforcement learning to optimize both the solution and the scaffold that guides the solution process. The repository emphasizes benchmark performance on Terminal-Bench, SWE-bench, NL2Repo, OpenClaw, and SWE Atlas while keeping the project MIT licensed and globally accessible.
    Downloads: 5 This Week
    Last Update:
    See Project
  • 3
    Kimi-K3

    Kimi-K3

    Open Frontier Intelligence

    Kimi K3 is an open-weight, native multimodal agentic model from Moonshot AI for advanced reasoning, coding, and knowledge work. It uses a 2.8-trillion-parameter mixture-of-experts architecture while activating 104 billion parameters per token. Its design combines Kimi Delta Attention, gated multihead latent attention, Attention Residuals, and Stable LatentMoE routing. The model can process text and images and supports a context window of more than one million tokens. It is built for long engineering sessions, large repository navigation, tool orchestration, research, visualization, and other extended tasks. ...
    Downloads: 13 This Week
    Last Update:
    See Project
  • 4
    OpenMythos

    OpenMythos

    A theoretical reconstruction of the Claude Mythos architecture

    ...It divides computation into three main stages, including a pre-processing phase, a looped recurrent reasoning block, and a final output refinement stage, creating a structured pipeline for inference. The architecture incorporates advanced techniques such as mixture-of-experts routing, adaptive computation time, and multiple attention mechanisms to dynamically allocate compute where needed. It is highly configurable through a centralized configuration system, allowing experimentation with different architectural parameters such as loop depth, attention type.
    Downloads: 15 This Week
    Last Update:
    See Project
  • Veeam Data Platform v13.1 Icon
    Veeam Data Platform v13.1

    Move workloads across hypervisors and clouds with no vendor lock-in. Try VDP free today.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try Now
  • 5
    vLLM Semantic Router

    vLLM Semantic Router

    System Level Intelligent Router for Mixture-of-Models at Cloud

    Semantic Router is an open-source system designed to intelligently route requests across multiple large language models based on the semantic meaning and complexity of user queries. Instead of sending every prompt to the same model, the system analyzes the intent and reasoning requirements of the request and dynamically selects the most appropriate model to process it. This approach allows developers to combine multiple models with different strengths, such as lightweight models for simple...
    Downloads: 22 This Week
    Last Update:
    See Project
  • 6
    Qwen3.5

    Qwen3.5

    Qwen3.5 is the large language model series developed by Qwen team

    ...Qwen3.5 builds on earlier Qwen generations by improving multilingual understanding, reasoning ability, and efficiency, while also introducing native multimodal capabilities that allow the model to work with both language and visual inputs. Architecturally, the system leverages modern large-scale training techniques and mixture-of-experts style efficiency so that very large parameter counts can be used while keeping inference practical.
    Downloads: 12 This Week
    Last Update:
    See Project
  • 7
    Wan2.2

    Wan2.2

    Wan2.2: Open and Advanced Large-Scale Video Generative Model

    Wan2.2 is a major upgrade to the Wan series of open and advanced large-scale video generative models, incorporating cutting-edge innovations to boost video generation quality and efficiency. It introduces a Mixture-of-Experts (MoE) architecture that splits the denoising process across specialized expert models, increasing total model capacity without raising computational costs. Wan2.2 integrates meticulously curated cinematic aesthetic data, enabling precise control over lighting, composition, color tone, and more, for high-quality, customizable video styles. ...
    Downloads: 125 This Week
    Last Update:
    See Project
  • 8
    LingBot-Video

    LingBot-Video

    Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

    LingBot-Video is a large-scale mixture-of-experts video generation model focused on embodied intelligence. It is designed to connect video synthesis with physical-world understanding instead of generating only visually appealing clips. The project includes dense and MoE model variants for text-to-image, text-to-video, and text-image-to-video workflows. Its training combines large-scale web video data with more than 70,000 hours of embodied data.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 9
    Flash-MoE

    Flash-MoE

    Running a big model on a small laptop

    Flash-MoE is a high-performance implementation of mixture-of-experts (MoE) architectures designed to optimize the efficiency and scalability of large AI models. It focuses on accelerating routing and computation by leveraging optimized kernels and memory management techniques, allowing models to dynamically select specialized sub-networks during inference. The project aims to reduce the computational cost typically associated with MoE systems while maintaining or improving performance. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 10
    Qwen3.6

    Qwen3.6

    Qwen3.6 is the large language model series developed by Qwen team

    ...One of its defining goals is to enhance “agentic coding,” enabling the model to reason across entire codebases, handle multi-step development tasks, and assist with complex software engineering workflows. The architecture incorporates modern techniques such as mixture-of-experts and hybrid attention mechanisms, allowing it to scale efficiently while maintaining strong performance.
    Downloads: 8 This Week
    Last Update:
    See Project
  • 11
    DeepSeek R1

    DeepSeek R1

    Open-source, high-performance AI model with advanced reasoning

    DeepSeek-R1 is an open-source large language model developed by DeepSeek, designed to excel in complex reasoning tasks across domains such as mathematics, coding, and language. DeepSeek R1 offers unrestricted access for both commercial and academic use. The model employs a Mixture of Experts (MoE) architecture, comprising 671 billion total parameters with 37 billion active parameters per token, and supports a context length of up to 128,000 tokens. DeepSeek-R1's training regimen uniquely integrates large-scale reinforcement learning (RL) without relying on supervised fine-tuning, enabling the model to develop advanced reasoning capabilities. ...
    Downloads: 117 This Week
    Last Update:
    See Project
  • 12
    Tile Kernels

    Tile Kernels

    A kernel library written in tilelang

    Tile Kernels is a DeepSeek kernel library written with TileLang for high-performance AI and machine-learning workloads. It contains specialized kernels for areas such as mixture-of-experts routing, quantization, batched transpose operations, Engram gating, and Manifold HyperConnection components. The project includes both optimized kernel implementations and PyTorch reference versions for comparison and validation. It is aimed at developers and researchers who work close to model internals and need efficient low-level building blocks. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    DeepSeekMath-V2

    DeepSeekMath-V2

    Towards self-verifiable mathematical reasoning

    ...Unlike general-purpose LLMs that might generate plausible-looking math but sometimes hallucinate or mishandle rigorous logic, Math-V2 is engineered to not only generate solutions but also self-verify them, meaning it examines the derivations, checks logical consistency, and flags or corrects mistakes, producing proofs + verification rather than just a final answer. Under the hood, Math-V2 uses a massive Mixture-of-Experts (MoE) architecture (activated parameter count reportedly in the hundreds of billions) derived from DeepSeek’s experimental base architecture. For math problems, it employs a generator-verifier loop: it first generates a candidate proof (or solution path), then runs a verifier that assesses correctness and completeness.
    Downloads: 15 This Week
    Last Update:
    See Project
  • 14
    MiMo-V2-Flash

    MiMo-V2-Flash

    MiMo-V2-Flash: Efficient Reasoning, Coding, and Agentic Foundation

    MiMo-V2-Flash is a large Mixture-of-Experts language model designed to deliver strong reasoning, coding, and agentic-task performance while keeping inference fast and cost-efficient. It uses an MoE setup where a very large total parameter count is available, but only a smaller subset is activated per token, which helps balance capability with runtime efficiency.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 15
    HunyuanImage-3.0

    HunyuanImage-3.0

    A Powerful Native Multimodal Model for Image Generation

    ...It unifies multimodal understanding and generation in a single autoregressive framework, combining text and image modalities seamlessly rather than relying on separate image-only diffusion components. It uses a Mixture-of-Experts (MoE) architecture with many expert subnetworks to scale efficiently, deploying only a subset of experts per token, which allows large parameter counts without linear inference cost explosion. The model is intended to be competitive with closed-source image generation systems, aiming for high fidelity, prompt adherence, fine detail, and even “world knowledge” reasoning (i.e. leveraging context, semantics, or common sense in generation). ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 16
    WASTE

    WASTE

    Run the full 2.78-trillion-parameter Kimi K3 model

    WASTE is an embeddable C inference engine for running extremely large mixture-of-experts models when the weights exceed available RAM. It keeps the shared model trunk in memory and streams only the experts selected for each token from fast NVMe storage. A bounded cache reuses recently needed experts, while lookahead routing begins disk reads before the next layer requires them. Its main target is the full 2.78-trillion-parameter Kimi K3 model, including multimodal image input. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    LLMs-Zero-to-Hero

    LLMs-Zero-to-Hero

    From nobody to big model (LLM) hero

    ...Rather than relying entirely on existing frameworks, the project encourages readers to implement important components themselves in order to gain a deeper understanding of how modern language models work internally. It includes explanations of dense transformer architectures, mixture-of-experts models, training pipelines, and techniques used in contemporary LLM development.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Awesome Stock Resources

    Awesome Stock Resources

    Collection of links for free stock photography, video and Illustration

    ...A collection of links to public domain photography resources. The photographs on some resources require Attribution unless otherwise stated on the website itself. These use a mixture of license, all of which have been linked to next to them. Some resources haven't specified any formal terms of use or licenses. A collection of illustration resources which contain a mixture of historical archive, contemporary and public domain assets. A collection of resources which contain stock graphical elements which don't fit in the other sections. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Qwen3-Coder

    Qwen3-Coder

    Qwen3-Coder is the code version of Qwen3

    Qwen3-Coder is the latest and most powerful agentic code model developed by the Qwen team at Alibaba Cloud. Its flagship version, Qwen3-Coder-480B-A35B-Instruct, features a massive 480 billion-parameter Mixture-of-Experts architecture with 35 billion active parameters, delivering top-tier performance on coding and agentic tasks. This model sets new state-of-the-art benchmarks among open models for agentic coding, browser-use, and tool-use, matching performance comparable to leading models like Claude Sonnet. Qwen3-Coder supports an exceptionally long context window of 256,000 tokens, extendable to 1 million tokens using Yarn, enabling repository-scale code understanding and generation. ...
    Downloads: 12 This Week
    Last Update:
    See Project
  • 20
    Xtuner

    Xtuner

    A Next-Generation Training Engine Built for Ultra-Large MoE Models

    Xtuner is a large-scale training engine designed for efficient training and fine-tuning of modern large language models, particularly mixture-of-experts architectures. The framework focuses on enabling scalable training for extremely large models while maintaining efficiency across distributed computing environments. Unlike traditional 3D parallel training strategies, XTuner introduces optimized parallelism techniques that simplify scaling and reduce system complexity when training massive models. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Ling-V2

    Ling-V2

    Ling-V2 is a MoE LLM provided and open-sourced by InclusionAI

    Ling-V2 is an open-source family of Mixture-of-Experts (MoE) large language models developed by the InclusionAI research organization with the goal of combining state-of-the-art performance, efficiency, and openness for next-generation AI applications. It introduces highly sparse architectures where only a fraction of the model’s parameters are activated per input token, enabling models like Ling-mini-2.0 to achieve reasoning and instruction-following capabilities on par with much larger dense models while remaining significantly more computationally efficient. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Ring

    Ring

    Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI

    Ring is a reasoning Mixture-of-Experts (MoE) large language model (LLM) developed by inclusionAI. It is built from or derived from Ling. Its design emphasizes reasoning, efficiency, and modular expert activation. In its “flash” variant (Ring-flash-2.0), it optimizes inference by activating only a subset of experts. It applies reinforcement learning/reasoning optimization techniques.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Heretic

    Heretic

    Fully automatic censorship removal for language models

    ...Designed for researchers and advanced users, Heretic makes it possible to study and experiment with uncensored model responses in a reproducible, automated way. The project can decensor many popular dense and some mixture-of-experts (MoE) models, supporting workflows that would otherwise require manual tuning. Beyond simple decensoring, Heretic includes research-oriented options for analyzing model internals and interpretability data.
    Downloads: 17 This Week
    Last Update:
    See Project
  • 24
    Wan2.1

    Wan2.1

    Wan2.1: Open and Advanced Large-Scale Video Generative Model

    ...The model supports text-to-video and image-to-video generation tasks with flexible resolution options suitable for various GPU hardware configurations. Wan2.1’s architecture balances generation quality and inference cost, paving the way for later improvements seen in Wan2.2 such as Mixture-of-Experts and enhanced aesthetics. It was trained on large-scale video and image datasets, providing generalization across diverse scenes and motion patterns.
    Downloads: 93 This Week
    Last Update:
    See Project
  • 25
    shimmy

    shimmy

    Python-free Rust inference server

    ...It supports modern model formats such as GGUF and SafeTensors and can automatically discover models stored locally or in common directories used by other AI tools. Advanced capabilities include CPU offloading for Mixture-of-Experts models and GPU acceleration, enabling large models to run on consumer hardware with limited VRAM.
    Downloads: 29 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • 3
  • 4
  • 5
  • Next