Showing 21 open source projects for "total text container"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • Veeam Data Platform v13.1 - Get Your Free Trial Icon
    Veeam Data Platform v13.1 - Get Your Free Trial

    Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try it Free
  • 1
    Wan2.2

    Wan2.2

    Wan2.2: Open and Advanced Large-Scale Video Generative Model

    ...Wan2.2 also open-sources a 5-billion parameter high-compression VAE-based hybrid text-image-to-video (TI2V) model that supports 720P video generation at 24fps on consumer-grade GPUs like the RTX 4090. It supports multiple video generation tasks including text-to-video.
    Downloads: 125 This Week
    Last Update:
    See Project
  • 2
    MiniMax-01

    MiniMax-01

    Large-language-model & vision-language-model based on Linear Attention

    MiniMax-01 is the official repository for two flagship models: MiniMax-Text-01, a long-context language model, and MiniMax-VL-01, a vision-language model built on top of it. MiniMax-Text-01 uses a hybrid attention architecture that blends Lightning Attention, standard softmax attention, and Mixture-of-Experts (MoE) routing to achieve both high throughput and long-context reasoning. It has 456 billion total parameters with 45.9 billion activated per token and is trained with advanced parallel strategies such as LASP+, varlen ring attention, and Expert Tensor Parallelism, enabling a training context of 1 million tokens and up to 4 million tokens at inference. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    VisualGLM-6B

    VisualGLM-6B

    Chinese and English multimodal conversational language model

    VisualGLM-6B is an open-source multimodal conversational language model developed by ZhipuAI that supports both images and text in Chinese and English. It builds on the ChatGLM-6B backbone, with 6.2 billion language parameters, and incorporates a BLIP2-Qformer visual module to connect vision and language. In total, the model has 7.8 billion parameters. Trained on a large bilingual dataset — including 30 million high-quality Chinese image-text pairs from CogView and 300 million English pairs — VisualGLM-6B is designed for image understanding, description, and question answering. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    Kimi K3

    Kimi K3

    Powerful, native multimodal AI agentic model

    Kimi K3 is an open-weight, multim is an open-weight, multimodal agentic AI model from Moonshot AI designed for advanced coding, research, reasoning, and knowledge work. Builtodal agentic AI model from Moonshot AI designed for advanced coding, research, reasoning, and knowledge work. Built with 2.8 trillion total parameters and a sparse mixture-of-experts architecture, it activates 104 billion parameters per token to with 2.8 trillion total parameters and a sparse mixture-of-experts architecture, it activates 104 billion parameters per token to deliver frontier-level performance more efficiently. Kimi K3 combines native text and image understanding deliver frontier-level performance more efficiently.
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    DiffusionGemma

    DiffusionGemma

    NVFP4 DiffusionGemma model for fast multimodal text generation

    DiffusionGemma 26B A4B IT NVFP4 is NVIDIA’s Model Optimizer quantized release of Google DeepMind’s DiffusionGemma 26B A4B IT model. It is an open-weights multimodal generative model that processes text, images, and video inputs to produce text output through discrete diffusion. Built on the Gemma 4 26B A4B Mixture-of-Experts architecture, it has 25.2B total parameters and 3.8B active parameters, balancing capability with efficient inference. Its diffusion-based generation produces tokens in parallel 256-token blocks, enabling very high-speed output, with reported generation above 1,100 tokens per second on NVIDIA Hopper H100 in FP8. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    Inkling-Small

    Inkling-Small

    Efficient multimodal MoE model for coding, tools, and reasoning

    Inkling-Small is an open-weight general-purpose multimodal model from Thinking Machines Lab, designed for agentic systems, coding assistants, chatbots, retrieval workflows, and natural-language applications. It accepts text, images, and audio as input and produces text output, with multilingual and multi-programming-language capabilities. The model uses a sparse Mixture-of-Experts architecture with 276B total parameters and 12B active per token, enabling strong performance with lower inference cost than a fully dense model of similar scale. Its 42-layer decoder routes each token through six of 256 specialized experts plus two shared experts, while hybrid local and global attention supports efficient processing. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Command A+

    Command A+

    4-bit Command A+ model for enterprise agents and multilingual tasks

    Command A+ 05-2026 W4A4 is a 4-bit quantized version of Cohere’s open-source Command A+ model, optimized for enterprise-grade agentic, multilingual, and reasoning-heavy workloads. It supports text and image inputs, generates text outputs, and uses a sparse Mixture-of-Experts Transformer architecture with 218B total parameters and 25B active parameters. The W4A4 release applies 4-bit weight and activation quantization mainly to MoE experts, preserving attention components at full precision to reduce quality loss while improving speed, latency, and hardware efficiency. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Qwen3.6-35B-A3B

    Qwen3.6-35B-A3B

    Open multimodal model for coding, agents, and long-context tasks

    ...Architecturally, it uses a Mixture-of-Experts design with 35B total parameters and 3B active, supports a native 262K-token context window, and can be extended to about 1M tokens with YaRN. It also performs strongly across coding, agent, vision, reasoning, and document-understanding benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    GLM-5.3-Flash

    GLM-5.3-Flash

    Efficient 320B multimodal MoE model for coding and autonomous agents

    ...GLM-5.3-Flash supports text and image inputs and is particularly optimized for coding and autonomous agent workloads, approaching larger frontier models on related benchmarks while improving over GLM-5.2. It supports local deployment through SGLang, vLLM, TokenSpeed, and KTransformers and is released under the MIT license.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Host LLMs in Production With On-Demand GPUs Icon
    Host LLMs in Production With On-Demand GPUs

    NVIDIA L4 GPUs. 5-second cold starts. Scale to zero when idle.

    Deploy your model, get an endpoint, pay only for compute time. No GPU provisioning or infrastructure management required.
    Start Free
  • 10
    MiMo-V2.5

    MiMo-V2.5

    Omnimodal AI model for agents, coding, and long-context tasks

    MiMo-V2.5 is a native omnimodal large language model developed by Xiaomi, designed for advanced agentic workflows, multimodal reasoning, and long-context processing. Built on a Mixture-of-Experts architecture with approximately 309B total parameters and around 15B activated per inference, it balances high capability with efficient execution. The model natively processes text, images, video, and audio within a unified system, enabling cross-modal understanding and complex task execution in a single pipeline. With a context window of up to 1 million tokens, it can handle large documents, extended conversations, and multi-step workflows without fragmentation. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Qwen3.6-35B-A3B-FP8

    Qwen3.6-35B-A3B-FP8

    FP8 Qwen model for efficient multimodal coding and agent tasks

    ...The model uses a Mixture-of-Experts design with 35B total parameters and 3B active, supports a native context window of 262,144 tokens, and can be extended to about 1,010,000 tokens with YaRN. It is compatible with major inference frameworks such as Transformers, vLLM, SGLang, and KTransformers, making it a practical high-performance option.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    MiMo-V2.6-Flash

    MiMo-V2.6-Flash

    Efficient 309B omnimodal MoE for coding, agents, vision, and audio

    MiMo-V2.6-Flash is Xiaomi MiMo’s efficiency-balanced open-weight omnimodal model, designed to scale reinforcement learning across coding, general agents, visual tasks, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 309B total parameters while activating only 15B per token, using 256 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and multi-session agent workflows. Training uses a unified mixed RL process rather than separate domain-specific runs, alongside asynchronous GRPO and groupwise agentic grading that rewards higher-quality and more efficient solutions. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    MiMo-V2.6-Pro

    MiMo-V2.6-Pro

    1T omnimodal MoE model for coding, agents, and long-horizon reasoning

    MiMo-V2.6-Pro is Xiaomi MiMo’s flagship open-weight omnimodal model, built to scale reinforcement learning toward self-improvement across coding, agents, vision, and cybersecurity. Its sparse Mixture-of-Experts architecture contains 1.02T total parameters with 42B activated per token, using 384 routed experts with eight active per token. The model natively processes text, images, video, and audio and supports a 1M-token context window for large repositories, extended tool traces, and multi-session agent workflows. MiMo-V2.6-Pro-RL uses a unified mixed reinforcement learning process rather than separate domain-specific runs, alongside groupwise agentic grading that rewards higher-quality and more efficient solutions. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    GLM-4.5-Air

    GLM-4.5-Air

    Compact hybrid reasoning language model for intelligent responses

    ...GLM-4.5-Air supports both English and Chinese, and is suitable for tasks involving text generation, coding, reasoning, and tool calling. Open-sourced under the MIT license, it is commercially usable and integrates with transformers, vLLM, and SGLang inference frameworks. It includes FP8 variants for faster inference and reduced memory requirements. Despite its smaller size compared to full GLM-4.5, GLM-4.5-Air maintains high performance.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Qwen3.8-2.4T-A95B

    Qwen3.8-2.4T-A95B

    Massive 2.4T MoE model for coding, agents, research, and reasoning

    ...Qwen3.8 also provides adjustable reasoning depth through low, medium, and xhigh reasoning-effort settings and preserves reasoning context across conversations. It is a text-only, thinking-first model and supports deployment through vLLM, SGLang, and TokenSpeed, with strong results across coding, tool use, research, and professional benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Nemotron 3.5 Lightning

    Nemotron 3.5 Lightning

    Efficient 30B MoE model for long-running agents and local inference

    ...NVIDIA also provides DSpark, Multi-Token Prediction, and DFlash speculative decoding methods to accelerate text generation. It supports English, coding languages, Spanish, French, German, Italian, and Japanese, and is intended for commercially deployable AI applications. The model can run on a single DGX Spark or H100 and integrates with inference frameworks including vLLM.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    Hy3

    Hy3

    Open code agent for Lean 4 proofs and formal software verification

    Leanstral 1.5 119B A6B is an open-source code agent model from Mistral AI designed specifically for Lean 4, a proof assistant used to express and verify complex mathematical objects and formal software specifications. Built as part of the Mistral Small 4 family, it combines multimodal capabilities with an efficient Mixture-of-Experts architecture containing 119B total parameters and 6.5B activated per token. The model uses 128 experts with four active for each token and supports a 256K-token context window, making it suitable for extended formal reasoning and large verification tasks. Leanstral accepts text and image inputs and produces text output, enabling multimodal workflows around mathematics, code, and specifications. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Leanstral 1.5

    Leanstral 1.5

    Open code agent for Lean 4 proofs and formal software verification

    Leanstral 1.5 119B A6B is an open-source code agent model from Mistral AI designed specifically for Lean 4, a proof assistant used to express and verify complex mathematical objects and formal software specifications. Built as part of the Mistral Small 4 family, it combines multimodal capabilities with an efficient Mixture-of-Experts architecture containing 119B total parameters and 6.5B activated per token. The model uses 128 experts with four active for each token and supports a 256K-token context window, making it suitable for extended formal reasoning and large verification tasks. Leanstral accepts text and image inputs and produces text output, enabling multimodal workflows around mathematics, code, and specifications. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Mistral Large 3 675B Instruct 2512 NVFP4

    Mistral Large 3 675B Instruct 2512 NVFP4

    Quantized 675B multimodal instruct model optimized for NVFP4

    ...The model integrates a 673B-parameter MoE language backbone with a 2.5B-parameter vision encoder, enabling rich multimodal analysis across text and images. Designed for efficient deployment, it runs on a single H100 or A100 node in NVFP4 while delivering performance similar to FP8 for short- and mid-context workloads.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Mistral Large 3 675B Instruct 2512

    Mistral Large 3 675B Instruct 2512

    Frontier-scale 675B multimodal instruct MoE model for enterprise AIMis

    Mistral Large 3 675B Instruct 2512 is a state-of-the-art multimodal granular Mixture-of-Experts model featuring 675B total parameters and 41B active parameters, trained from scratch on 3,000 H200 GPUs. As the instruct-tuned FP8 variant, it is optimized for reliable instruction following, agentic workflows, production-grade assistants, and long-context enterprise tasks. It incorporates a massive 673B-parameter language MoE backbone and a 2.5B-parameter vision encoder, enabling rich multimodal understanding across text and images. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Inkling

    Inkling

    Frontier multimodal MoE model for coding and AI agent workflows

    Inkling is Thinking Machines Lab’s first open-weight flagship multimodal Mixture-of-Experts model, designed for advanced reasoning, coding, and autonomous agent workflows. It contains 975B total parameters with 41B active parameters per token, balancing frontier-level capability with efficient sparse inference. The model natively processes text, images, audio, and video within a unified architecture and supports an exceptionally large 1 million token context window for long-document reasoning, repository-scale coding, and agentic execution. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • Next