Search Results for "bayesian mixture model" - Page 7

Showing 170 open source projects for "bayesian mixture model"

View related business solutions
  • Veeam Data Platform v13.1 Icon
    Veeam Data Platform v13.1

    Move workloads across hypervisors and clouds with no vendor lock-in. Try VDP free today.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try Now
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    Mistral Large 3 675B Base 2512

    Mistral Large 3 675B Base 2512

    Frontier-scale 675B multimodal base model for custom AI training

    Mistral Large 3 675B Base 2512 is the foundational, pre-trained version of the Mistral Large 3 family, built as a frontier-scale multimodal Mixture-of-Experts model with 41B active parameters and a total size of 675B. It is trained from scratch using 3000 H200 GPUs, making it one of the most advanced and compute-intensive open-weight models available. As the base version, it is not fine-tuned for instruction following or reasoning, making it ideal for teams planning their own domain-specific finetuning or custom training pipelines. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Inkling

    Inkling

    Frontier multimodal MoE model for coding and AI agent workflows

    Inkling is Thinking Machines Lab’s first open-weight flagship multimodal Mixture-of-Experts model, designed for advanced reasoning, coding, and autonomous agent workflows. It contains 975B total parameters with 41B active parameters per token, balancing frontier-level capability with efficient sparse inference. The model natively processes text, images, audio, and video within a unified architecture and supports an exceptionally large 1 million token context window for long-document reasoning, repository-scale coding, and agentic execution. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Command A+

    Command A+

    4-bit Command A+ model for enterprise agents and multilingual tasks

    Command A+ 05-2026 W4A4 is a 4-bit quantized version of Cohere’s open-source Command A+ model, optimized for enterprise-grade agentic, multilingual, and reasoning-heavy workloads. It supports text and image inputs, generates text outputs, and uses a sparse Mixture-of-Experts Transformer architecture with 218B total parameters and 25B active parameters. The W4A4 release applies 4-bit weight and activation quantization mainly to MoE experts, preserving attention components at full precision to reduce quality loss while improving speed, latency, and hardware efficiency. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    MiMo-V2.5-Pro

    MiMo-V2.5-Pro

    Flagship MoE model for long-context agents and complex coding

    MiMo-V2.5-Pro is Xiaomi’s flagship Mixture-of-Experts (MoE) model built for the most demanding agentic, software engineering, and long-horizon reasoning tasks. It features approximately 1.02 trillion total parameters with 42B activated per inference, balancing extreme capability with efficient execution. The model supports a 1 million token context window, enabling it to maintain coherence across long workflows involving thousands of tool calls and multi-step reasoning chains. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Earn up to 16% annual interest with Nexo. Icon
    Earn up to 16% annual interest with Nexo.

    More flexibility. More control.

    Generate interest, access liquidity without selling, and execute trades seamlessly. All in one platform. Geographic restrictions, eligibility, and terms apply.
    Get started with Nexo.
  • 5
    MiMo-V2.5

    MiMo-V2.5

    Omnimodal AI model for agents, coding, and long-context tasks

    MiMo-V2.5 is a native omnimodal large language model developed by Xiaomi, designed for advanced agentic workflows, multimodal reasoning, and long-context processing. Built on a Mixture-of-Experts architecture with approximately 309B total parameters and around 15B activated per inference, it balances high capability with efficient execution. The model natively processes text, images, video, and audio within a unified system, enabling cross-modal understanding and complex task execution in a single pipeline. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    GigaChat 3 Ultra

    GigaChat 3 Ultra

    High-performance MoE model with MLA, MTP, and multilingual reasoning

    GigaChat 3 Ultra is a flagship instruct-model built on a custom Mixture-of-Experts architecture with 702B total and 36B active parameters. It leverages Multi-head Latent Attention to compress the KV cache into latent vectors, dramatically reducing memory demand and improving inference speed at scale. The model also employs Multi-Token Prediction, enabling multi-step token generation in a single pass for up to 40% faster output through speculative and parallel decoding techniques. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Qwen3-Next

    Qwen3-Next

    Qwen3-Next: 80B instruct LLM with ultra-long context up to 1M tokens

    Qwen3-Next-80B-A3B-Instruct is the flagship release in the Qwen3-Next series, designed as a next-generation foundation model for ultra-long context and efficient reasoning. With 80B total parameters and 3B activated at a time, it leverages hybrid attention (Gated DeltaNet + Gated Attention) and a high-sparsity Mixture-of-Experts architecture to achieve exceptional efficiency. The model natively supports a context length of 262K tokens and can be extended up to 1 million tokens using RoPE scaling (YaRN), making it highly capable for processing large documents and extended conversations. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Mistral Large 3 675B Instruct 2512

    Mistral Large 3 675B Instruct 2512

    Frontier-scale 675B multimodal instruct MoE model for enterprise AIMis

    Mistral Large 3 675B Instruct 2512 is a state-of-the-art multimodal granular Mixture-of-Experts model featuring 675B total parameters and 41B active parameters, trained from scratch on 3,000 H200 GPUs. As the instruct-tuned FP8 variant, it is optimized for reliable instruction following, agentic workflows, production-grade assistants, and long-context enterprise tasks. It incorporates a massive 673B-parameter language MoE backbone and a 2.5B-parameter vision encoder, enabling rich multimodal understanding across text and images. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Kimi K2.6

    Kimi K2.6

    Multimodal agent model for coding, orchestration, and autonomy

    ...One of its most distinctive capabilities is horizontal agent scaling, supporting up to 300 sub-agents and 4,000 coordinated steps in a single run, which enables parallel task decomposition and end-to-end completion of outputs such as documents, websites, and spreadsheets. Architecturally, it uses a 1T-parameter Mixture-of-Experts design with 32B activated parameters, a MoonViT vision encoder, and a 256K context window.
    Downloads: 0 This Week
    Last Update:
    See Project
  • AI-generated apps that pass security review Icon
    AI-generated apps that pass security review

    Stop waiting on engineering. Build production-ready internal tools with AI—on your company data, in your cloud.

    Retool lets you generate dashboards, admin panels, and workflows directly on your data. Type something like “Build me a revenue dashboard on my Stripe data” and get a working app with security, permissions, and compliance built in from day one. Whether on our cloud or self-hosted, create the internal software your team needs without compromising enterprise standards or control.
    Try Retool free
  • 10
    Smaug Flash

    Smaug Flash

    304B MoE model optimized for agentic coding and long-running workflows

    Smaug Flash is an open-weight agentic coding model from Abacus.AI, fine-tuned from DeepSeek-V4-Flash-0731 to improve autonomous software engineering, tool use, and long-running agent workflows. It uses a 304B-parameter Mixture-of-Experts architecture with 43 layers, 256 routed experts, six selected experts per token, and one shared expert. Its attention system combines Multi-Head Latent Attention with a sparse token indexer, while DSpark multi-token prediction provides speculative decoding for faster generation. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Hy3

    Hy3

    Open code agent for Lean 4 proofs and formal software verification

    Leanstral 1.5 119B A6B is an open-source code agent model from Mistral AI designed specifically for Lean 4, a proof assistant used to express and verify complex mathematical objects and formal software specifications. Built as part of the Mistral Small 4 family, it combines multimodal capabilities with an efficient Mixture-of-Experts architecture containing 119B total parameters and 6.5B activated per token.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Leanstral 1.5

    Leanstral 1.5

    Open code agent for Lean 4 proofs and formal software verification

    Leanstral 1.5 119B A6B is an open-source code agent model from Mistral AI designed specifically for Lean 4, a proof assistant used to express and verify complex mathematical objects and formal software specifications. Built as part of the Mistral Small 4 family, it combines multimodal capabilities with an efficient Mixture-of-Experts architecture containing 119B total parameters and 6.5B activated per token.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    DiffusionGemma

    DiffusionGemma

    NVFP4 DiffusionGemma model for fast multimodal text generation

    DiffusionGemma 26B A4B IT NVFP4 is NVIDIA’s Model Optimizer quantized release of Google DeepMind’s DiffusionGemma 26B A4B IT model. It is an open-weights multimodal generative model that processes text, images, and video inputs to produce text output through discrete diffusion. Built on the Gemma 4 26B A4B Mixture-of-Experts architecture, it has 25.2B total parameters and 3.8B active parameters, balancing capability with efficient inference.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Mistral Large 3 675B Instruct 2512 NVFP4

    Mistral Large 3 675B Instruct 2512 NVFP4

    Quantized 675B multimodal instruct model optimized for NVFP4

    Mistral Large 3 675B Instruct 2512 NVFP4 is a frontier-scale multimodal Mixture-of-Experts model featuring 675B total parameters and 41B active parameters, trained from scratch on 3,000 H200 GPUs. This NVFP4 checkpoint is a post-training-activation quantized version of the original instruct model, created through a collaboration between Mistral AI, vLLM, and Red Hat using llm-compressor. It retains the same instruction-tuned behavior as the FP8 model, making it ideal for production assistants, agentic workflows, scientific tasks, and long-context enterprise systems. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Ling 3.0 Tiny

    Ling 3.0 Tiny

    Lightweight MoE model for local reasoning, coding, and AI agents

    Ling-3.0-tiny is inclusionAI’s lightweight hybrid reasoning Mixture-of-Experts model, designed to provide capable reasoning and agentic performance at low inference cost. It contains 7.9B total parameters while activating only 1.3B per token, using a hybrid architecture that alternates Kimi Delta Attention and Multi-Head Latent Attention with a sparse 128-expert MoE. The model supports both fast responses and configurable multi-step thinking, covering general agents, coding, mathematics, scientific reasoning, and instruction following. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Nemotron 3.5 Lightning

    Nemotron 3.5 Lightning

    Efficient 30B MoE model for long-running agents and local inference

    NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4 is an open large language model optimized for efficient autonomous agents, sub-agent deployments, and local inference. It uses a hybrid Mixture-of-Experts architecture combining Mamba-2, MoE, and selected attention layers, with 30B total parameters but only 3B active during inference. The model supports context windows up to 1 million tokens, enabling long-running workflows and large-context reasoning.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    GLM-4.5-Air

    GLM-4.5-Air

    Compact hybrid reasoning language model for intelligent responses

    GLM-4.5-Air is a multilingual large language model with 106 billion total parameters and 12 billion active parameters, designed for conversational AI and intelligent agents. It is part of the GLM-4.5 family developed by Zhipu AI, offering hybrid reasoning capabilities via two modes: a thinking mode for complex reasoning and tool use, and a non-thinking mode for immediate responses. The model is optimized for efficiency and deployment, delivering strong results across 12 industry benchmarks,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Kimi K2.7 Code

    Kimi K2.7 Code

    Coding-focused Kimi model for long-horizon agent workflows

    Kimi K2.7 Code is a coding-focused agentic model built on Kimi K2.6, designed for long-horizon software engineering, autonomous coding workflows, and complex tool-based execution. It improves end-to-end task completion across real-world programming scenarios while reducing thinking-token usage by about 30% compared with K2.6. Architecturally, it uses a 1T-parameter Mixture-of-Experts design with 32B activated parameters, 61 layers, 384 experts, a 256K-token context window, and a MoonViT vision encoder. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Hy4 preview

    Hy4 preview

    770B MoE model for coding, research, reasoning, and long-context work

    Hy4 preview is Tencent’s open-weight flagship Mixture-of-Experts language model designed for advanced reasoning, software engineering, productivity, scientific research, and long-horizon tasks. It contains 770B backbone parameters while activating 49B per token across 78 layers, with 256 routed experts and one shared expert in each MoE layer. Its architecture uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse-index reuse and identity Hyper-Connections to improve information flow. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Solar Open 2

    Solar Open 2

    Efficient 250B MoE model for agents, coding, and long-context work

    Solar Open 2 is Upstage’s 250B-A15B open-weight large language model designed for agentic workflows, office productivity, document-intensive tasks, coding, and reasoning. Its Hybrid-Attention Mixture-of-Experts architecture contains 250B total parameters while activating only 15B per token, combining three linear-attention layers with one softmax-attention layer for efficient inference. The model supports a native 1M-token context window and uses NoPE instead of rotary positional encoding, reducing long-context KV-cache requirements. ...
    Downloads: 0 This Week
    Last Update:
    See Project