Showing 1921 open source projects for "context"

View related business solutions
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • Cut Data Warehouse Costs by 54% Icon
    Cut Data Warehouse Costs by 54%

    Easily migrate from Snowflake, Redshift, or Databricks with free tools.

    BigQuery delivers 54% lower TCO with exabyte scale and flexible pricing. Free migration tools handle the SQL translation automatically.
    Try Free
  • 1
    Qwen3.6-35B-A3B

    Qwen3.6-35B-A3B

    Open multimodal model for coding, agents, and long-context tasks

    ...A notable addition is thinking preservation, which allows the model to retain reasoning context from earlier messages, improving iterative work and reducing redundant computation. Architecturally, it uses a Mixture-of-Experts design with 35B total parameters and 3B active, supports a native 262K-token context window, and can be extended to about 1M tokens with YaRN. It also performs strongly across coding, agent, vision, reasoning, and document-understanding benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Qwen3.6-27B

    Qwen3.6-27B

    Dense multimodal Qwen model for coding, agents, and long context

    ...It also introduces thinking preservation, allowing it to retain reasoning traces from earlier turns to improve consistency, reduce repeated computation, and support iterative agent workflows. Qwen3.6-27B natively supports a 262K-token context window and can be extended to about 1M tokens with YaRN for ultra-long tasks. It is compatible with Transformers, vLLM, SGLang, and KTransformers, supports tool calling through Qwen-Agent.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Qwen3.8-2.4T-A95B

    Qwen3.8-2.4T-A95B

    Massive 2.4T MoE model for coding, agents, research, and reasoning

    ...The model emphasizes reliable autonomous execution, including stronger planning, environment feedback handling, and end-to-end completion of complex multi-step workflows. It natively supports a 262K-token context window that can be extended beyond one million tokens. Qwen3.8 also provides adjustable reasoning depth through low, medium, and xhigh reasoning-effort settings and preserves reasoning context across conversations. It is a text-only, thinking-first model and supports deployment through vLLM, SGLang, and TokenSpeed, with strong results across coding, tool use, research, and professional benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    MiMo-V2.5-Pro

    MiMo-V2.5-Pro

    Flagship MoE model for long-context agents and complex coding

    ...It features approximately 1.02 trillion total parameters with 42B activated per inference, balancing extreme capability with efficient execution. The model supports a 1 million token context window, enabling it to maintain coherence across long workflows involving thousands of tool calls and multi-step reasoning chains. Architecturally, it uses a hybrid attention system combining Sliding Window Attention and Global Attention to significantly reduce memory usage while preserving long-context performance. It also integrates multi-token prediction modules that accelerate inference and improve reinforcement learning efficiency. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Get Started Free
  • 5
    MiMo-V2.5

    MiMo-V2.5

    Omnimodal AI model for agents, coding, and long-context tasks

    MiMo-V2.5 is a native omnimodal large language model developed by Xiaomi, designed for advanced agentic workflows, multimodal reasoning, and long-context processing. Built on a Mixture-of-Experts architecture with approximately 309B total parameters and around 15B activated per inference, it balances high capability with efficient execution. The model natively processes text, images, video, and audio within a unified system, enabling cross-modal understanding and complex task execution in a single pipeline. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    DeepSeek-V4-Flash

    DeepSeek-V4-Flash

    Efficient MoE model for million-token reasoning and coding

    DeepSeek-V4-Flash is a preview Mixture-of-Experts language model built for efficient million-token context intelligence. It has 284B total parameters with 13B activated and supports a 1M-token context window, making it suitable for long-document reasoning, complex coding, agentic workflows, and large-scale information processing. The model uses a hybrid attention architecture that combines Compressed Sparse Attention and Heavily Compressed Attention to improve long-context efficiency, while Manifold-Constrained Hyper-Connections strengthen signal stability across layers. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Qwen3-Next

    Qwen3-Next

    Qwen3-Next: 80B instruct LLM with ultra-long context up to 1M tokens

    Qwen3-Next-80B-A3B-Instruct is the flagship release in the Qwen3-Next series, designed as a next-generation foundation model for ultra-long context and efficient reasoning. With 80B total parameters and 3B activated at a time, it leverages hybrid attention (Gated DeltaNet + Gated Attention) and a high-sparsity Mixture-of-Experts architecture to achieve exceptional efficiency. The model natively supports a context length of 262K tokens and can be extended up to 1 million tokens using RoPE scaling (YaRN), making it highly capable for processing large documents and extended conversations. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    Hunyuan-A13B-Instruct

    Hunyuan-A13B-Instruct

    Efficient 13B MoE language model with long context and reasoning modes

    ...Open-source under a custom license, it's ideal for researchers and developers seeking scalable, high-context AI capabilities with optimized inference.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Postproxy-MCP

    Postproxy-MCP

    MCP (Model Context Protocol) server for integrating PostProxy API

    PostProxy MCP is a Model Context Protocol (MCP) server that integrates the PostProxy API directly into Claude Code, enabling AI-assisted publishing to social media platforms like Instagram, YouTube, TikTok, Facebook, LinkedIn, X/Twitter, and Threads.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Host LLMs in Production With On-Demand GPUs Icon
    Host LLMs in Production With On-Demand GPUs

    NVIDIA L4 GPUs. 5-second cold starts. Scale to zero when idle.

    Deploy your model, get an endpoint, pay only for compute time. No GPU provisioning or infrastructure management required.
    Try Free
  • 10
    Hy3 preview

    Hy3 preview

    Efficient MoE model for reasoning, coding, and AI agent workflows

    Hy3 preview is Tencent Hunyuan’s latest open-weight Mixture-of-Experts language model, designed for advanced reasoning, coding, instruction following, and autonomous agent workflows. It is the first model built on Tencent’s rebuilt training infrastructure and introduces significant improvements in context learning, software engineering, and tool-based task execution. The model features 295B total parameters with only 21B activated during inference, plus a dedicated 3.8B Multi-Token Prediction (MTP) layer that accelerates generation through speculative decoding. Architecturally, it uses 192 routed experts with top-8 activation, a dense-MoE hybrid design, and a native 256K-token context window. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Qwen2-7B-Instruct

    Qwen2-7B-Instruct

    Instruction-tuned 7B language model for chat and complex tasks

    ...Built on a transformer architecture with SwiGLU activation and group query attention, it is optimized for chat, reasoning, coding, multilingual tasks, and extended context understanding up to 131,072 tokens. The model was pretrained on a large-scale dataset and aligned via supervised fine-tuning and direct preference optimization. It shows strong performance across benchmarks such as MMLU, MT-Bench, GSM8K, and Humaneval, often surpassing similarly sized open-source models. Designed for conversational use, it integrates with Hugging Face Transformers and supports long-context applications via YARN and vLLM for efficient deployment.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    QwQ-32B

    QwQ-32B

    QwQ-32B is a reasoning-focused language model for complex tasks

    ...For long input handling, it supports YaRN (Yet another RoPE Namer) for context scaling.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    LongCat-2.0

    LongCat-2.0

    Trillion-parameter MoE model for coding and million-token reasoning

    ...The model was pretrained on more than 35 trillion tokens and trained entirely on a large-scale cluster of domestically developed AI accelerators, demonstrating stable frontier-scale training without rollback events. LongCat-2.0 introduces LongCat Sparse Attention and extensive 1M-context training, enabling native processing of million-token inputs for long-document analysis, repository-scale coding, and complex multi-step reasoning. Dedicated post-training further strengthens coding and agent performance, producing competitive benchmark results against leading proprietary models.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Devstral Small 2

    Devstral Small 2

    Lightweight 24B agentic coding model with vision and long context

    ...The model achieves competitive performance on SWE-bench, validating its effectiveness for real-world coding and automation tasks. It introduces vision capabilities, enabling image understanding alongside text for more versatile development workflows. Devstral Small 2 supports a 256k context window, allowing it to reason across large repositories, long diffs, and extended technical contexts. Its architecture improves generalization across diverse prompts and coding environments while leveraging advanced attention scaling techniques.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Qwen2.5-14B-Instruct

    Qwen2.5-14B-Instruct

    Powerful 14B LLM with strong instruction and long-text handling

    Qwen2.5-14B-Instruct is a powerful instruction-tuned language model developed by the Qwen team, based on the Qwen2.5 architecture. It features 14.7 billion parameters and is optimized for tasks like dialogue, long-form generation, and structured output. The model supports context lengths up to 128K tokens and can generate up to 8K tokens, making it suitable for long-context applications. It demonstrates improved performance in coding, mathematics, and multilingual understanding across over 29 languages. Qwen2.5-14B-Instruct is built on a transformer backbone with RoPE, SwiGLU, RMSNorm, and attention QKV bias. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Nemotron 3.5 Lightning

    Nemotron 3.5 Lightning

    Efficient 30B MoE model for long-running agents and local inference

    ...It uses a hybrid Mixture-of-Experts architecture combining Mamba-2, MoE, and selected attention layers, with 30B total parameters but only 3B active during inference. The model supports context windows up to 1 million tokens, enabling long-running workflows and large-context reasoning. Its NVFP4 quantization reduces deployment requirements while targeting NVIDIA hardware ranging from DGX Spark and RTX 5090 systems to H100, H200, and GB200 accelerators. NVIDIA also provides DSpark, Multi-Token Prediction, and DFlash speculative decoding methods to accelerate text generation. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    DiffusionGemma

    DiffusionGemma

    NVFP4 DiffusionGemma model for fast multimodal text generation

    ...Its diffusion-based generation produces tokens in parallel 256-token blocks, enabling very high-speed output, with reported generation above 1,100 tokens per second on NVIDIA Hopper H100 in FP8. The model supports a 256K-token context window, configurable thinking mode, native function calling, structured JSON output, and multilingual inference across 35+ languages. The NVFP4 quantization reduces weights and activations from 16-bit to 4-bit, lowering disk size and GPU memory needs for vLLM deployment.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Gemma 4 12B

    Gemma 4 12B

    Unified multimodal Gemma model for local coding and reasoning

    ...It supports text, image, audio, and video inputs with text output, making it useful for transcription, image understanding, video analysis, coding, and agentic workflows. The model has 11.95B parameters, 48 layers, a 256K-token context window, and support for over 140 languages. It also includes configurable thinking modes, native system prompt support, function calling, and strong benchmark performance for its size. It is optimized for consumer GPUs, workstations, and streamlined local deployment.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Qwen3.6-35B-A3B-FP8

    Qwen3.6-35B-A3B-FP8

    FP8 Qwen model for efficient multimodal coding and agent tasks

    ...It is a multimodal open-weight model that combines a causal language model with a vision encoder, supporting text, image, and video inputs. Built for stability and real-world developer use, it emphasizes agentic coding, repository-level reasoning, and productive long-context workflows. A key capability is thinking preservation, which allows the model to retain reasoning traces from earlier messages, helping reduce repeated computation and improving consistency in iterative tasks. The model uses a Mixture-of-Experts design with 35B total parameters and 3B active, supports a native context window of 262,144 tokens, and can be extended to about 1,010,000 tokens with YaRN. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Kimi K2.7 Code

    Kimi K2.7 Code

    Coding-focused Kimi model for long-horizon agent workflows

    ...It improves end-to-end task completion across real-world programming scenarios while reducing thinking-token usage by about 30% compared with K2.6. Architecturally, it uses a 1T-parameter Mixture-of-Experts design with 32B activated parameters, 61 layers, 384 experts, a 256K-token context window, and a MoonViT vision encoder. The model supports image and video input, native INT4 quantization, interleaved thinking, and multi-step tool calling. It also forces preserve-thinking mode by default, retaining full reasoning context across multi-turn interactions to improve coding-agent consistency. K2.7 Code is recommended for use through Kimi Code CLI and can be deployed with vLLM, SGLang, or KTransformers.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Mistral Large 3 675B Base 2512

    Mistral Large 3 675B Base 2512

    Frontier-scale 675B multimodal base model for custom AI training

    ...As the base version, it is not fine-tuned for instruction following or reasoning, making it ideal for teams planning their own domain-specific finetuning or custom training pipelines. The model is engineered for reliability, long-context comprehension, and stable performance across many enterprise, scientific, and knowledge-intensive workloads. Its architecture includes a powerful language MoE and a 2.5B-parameter vision encoder, enabling multimodal understanding out of the box. Mistral Large 3 Base supports deployment on-premises using FP8 or NVFP4 formats, enabling high-performance workflows on B200, H200, H100, or A100 hardware.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Mistral Large 3 675B Instruct 2512 NVFP4

    Mistral Large 3 675B Instruct 2512 NVFP4

    Quantized 675B multimodal instruct model optimized for NVFP4

    ...This NVFP4 checkpoint is a post-training-activation quantized version of the original instruct model, created through a collaboration between Mistral AI, vLLM, and Red Hat using llm-compressor. It retains the same instruction-tuned behavior as the FP8 model, making it ideal for production assistants, agentic workflows, scientific tasks, and long-context enterprise systems. The model integrates a 673B-parameter MoE language backbone with a 2.5B-parameter vision encoder, enabling rich multimodal analysis across text and images. Designed for efficient deployment, it runs on a single H100 or A100 node in NVFP4 while delivering performance similar to FP8 for short- and mid-context workloads.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Mistral Large 3 675B Instruct 2512

    Mistral Large 3 675B Instruct 2512

    Frontier-scale 675B multimodal instruct MoE model for enterprise AIMis

    ...With a 256k context window, it excels at long-document comprehension, deep retrieval workflows, and complex knowledge-intensive tasks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Ministral 3 3B Reasoning 2512

    Ministral 3 3B Reasoning 2512

    Compact 3B-param multimodal model for efficient on-device reasoning

    ...It supports dozens of languages, allowing it to function across global and multilingual contexts. The model retains strong system-prompt adherence, supports function-calling with structured JSON output, and offers a large 256k token context window for extended context reasoning.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    This task library is built upon libspe2 from IBM Cell SDK. It gives a clean task interface to Cell programmers to start jobs in SPUs by hiding the tedious context/pthread creation, mailbox/signal/interrupt mailbox communication, etc.
    Downloads: 0 This Week
    Last Update:
    See Project