Showing 2140 open source projects for "context"

View related business solutions
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Start Free
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • 1
    LongCat-2.0

    LongCat-2.0

    Trillion-parameter MoE model for coding and million-token reasoning

    ...The model was pretrained on more than 35 trillion tokens and trained entirely on a large-scale cluster of domestically developed AI accelerators, demonstrating stable frontier-scale training without rollback events. LongCat-2.0 introduces LongCat Sparse Attention and extensive 1M-context training, enabling native processing of million-token inputs for long-document analysis, repository-scale coding, and complex multi-step reasoning. Dedicated post-training further strengthens coding and agent performance, producing competitive benchmark results against leading proprietary models.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    Devstral Small 2

    Devstral Small 2

    Lightweight 24B agentic coding model with vision and long context

    ...The model achieves competitive performance on SWE-bench, validating its effectiveness for real-world coding and automation tasks. It introduces vision capabilities, enabling image understanding alongside text for more versatile development workflows. Devstral Small 2 supports a 256k context window, allowing it to reason across large repositories, long diffs, and extended technical contexts. Its architecture improves generalization across diverse prompts and coding environments while leveraging advanced attention scaling techniques.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Qwen2.5-14B-Instruct

    Qwen2.5-14B-Instruct

    Powerful 14B LLM with strong instruction and long-text handling

    Qwen2.5-14B-Instruct is a powerful instruction-tuned language model developed by the Qwen team, based on the Qwen2.5 architecture. It features 14.7 billion parameters and is optimized for tasks like dialogue, long-form generation, and structured output. The model supports context lengths up to 128K tokens and can generate up to 8K tokens, making it suitable for long-context applications. It demonstrates improved performance in coding, mathematics, and multilingual understanding across over 29 languages. Qwen2.5-14B-Instruct is built on a transformer backbone with RoPE, SwiGLU, RMSNorm, and attention QKV bias. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    Hy4 preview

    Hy4 preview

    770B MoE model for coding, research, reasoning, and long-context work

    ...Its architecture uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse-index reuse and identity Hyper-Connections to improve information flow. A native 10B-parameter Multi-Token Prediction layer enables speculative decoding for faster inference. Hy4-preview supports a native 1M-token context window, allowing it to process large codebases, numerous files, and complex extended workflows. Tencent specifically optimized the model for software engineering, office and financial analysis, game development, and scientific research.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Paessler: Easy to Use With Enterprise Power. Free Trial Icon
    Paessler: Easy to Use With Enterprise Power. Free Trial

    A low-code dashboard makes monitoring intuitive for any admin, while scripting and custom sensors give experts full control.

    You shouldn't have to choose between a monitoring tool that's easy to use and one that's powerful enough for a complex environment. PRTG's low-code interface lets any admin build dashboards, set alerts and monitor devices without scripting, while custom sensors and full API access are there when your team needs deeper control. One platform, no compromise. Download a free 30-day trial now.
    Get Free Download
  • 5
    Qwen3.8-Flash-Next

    Qwen3.8-Flash-Next

    Efficient multimodal MoE model for coding, reasoning, and AI agents

    ...Its hybrid architecture combines Gated DeltaNet with Qwen Sparse Attention (QSA), which processes micro-blocks rather than individual tokens to reduce latency in long-context agent workloads. The model also introduces Gated Residual connections and scalable n-gram embeddings to improve efficiency while limiting inference overhead. It contains 512 MoE experts, activating 10 routed experts plus one shared expert per token. Qwen3.8-Flash-Next natively handles text, images, and video and supports a 262K-token context window extensible to 1M tokens. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    GLM-5.3-Flash

    GLM-5.3-Flash

    Efficient 320B multimodal MoE model for coding and autonomous agents

    GLM-5.3-Flash is Z.ai’s natively multimodal model designed for efficient coding, agentic engineering, reasoning, and long-context workloads. It uses a sparse architecture with 320B total parameters and only 18B active parameters, targeting high capability with substantially lower inference costs. The model introduces a hybrid architecture combining sparse and linear attention to reduce long-context serving costs while retaining precise understanding across large inputs. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Nemotron 3.5 Lightning

    Nemotron 3.5 Lightning

    Efficient 30B MoE model for long-running agents and local inference

    ...It uses a hybrid Mixture-of-Experts architecture combining Mamba-2, MoE, and selected attention layers, with 30B total parameters but only 3B active during inference. The model supports context windows up to 1 million tokens, enabling long-running workflows and large-context reasoning. Its NVFP4 quantization reduces deployment requirements while targeting NVIDIA hardware ranging from DGX Spark and RTX 5090 systems to H100, H200, and GB200 accelerators. NVIDIA also provides DSpark, Multi-Token Prediction, and DFlash speculative decoding methods to accelerate text generation. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    DiffusionGemma

    DiffusionGemma

    NVFP4 DiffusionGemma model for fast multimodal text generation

    ...Its diffusion-based generation produces tokens in parallel 256-token blocks, enabling very high-speed output, with reported generation above 1,100 tokens per second on NVIDIA Hopper H100 in FP8. The model supports a 256K-token context window, configurable thinking mode, native function calling, structured JSON output, and multilingual inference across 35+ languages. The NVFP4 quantization reduces weights and activations from 16-bit to 4-bit, lowering disk size and GPU memory needs for vLLM deployment.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Gemma 4 12B

    Gemma 4 12B

    Unified multimodal Gemma model for local coding and reasoning

    ...It supports text, image, audio, and video inputs with text output, making it useful for transcription, image understanding, video analysis, coding, and agentic workflows. The model has 11.95B parameters, 48 layers, a 256K-token context window, and support for over 140 languages. It also includes configurable thinking modes, native system prompt support, function calling, and strong benchmark performance for its size. It is optimized for consumer GPUs, workstations, and streamlined local deployment.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 99.99% Uptime for MySQL and PostgreSQL Databases Icon
    99.99% Uptime for MySQL and PostgreSQL Databases

    Sub-second maintenance. 2x read/write performance. Built-in vector search for AI apps.

    Cloud SQL Enterprise Plus delivers near-zero downtime with 35 days of point-in-time recovery. Supports MySQL, PostgreSQL, and SQL Server.
    Start Free
  • 10
    Qwen3.6-35B-A3B-FP8

    Qwen3.6-35B-A3B-FP8

    FP8 Qwen model for efficient multimodal coding and agent tasks

    ...It is a multimodal open-weight model that combines a causal language model with a vision encoder, supporting text, image, and video inputs. Built for stability and real-world developer use, it emphasizes agentic coding, repository-level reasoning, and productive long-context workflows. A key capability is thinking preservation, which allows the model to retain reasoning traces from earlier messages, helping reduce repeated computation and improving consistency in iterative tasks. The model uses a Mixture-of-Experts design with 35B total parameters and 3B active, supports a native context window of 262,144 tokens, and can be extended to about 1,010,000 tokens with YaRN. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Kolibri-1

    Kolibri-1

    Efficient bilingual MoE model for reasoning, coding, RAG, and agents

    ...Kolibri supports explicit reasoning with configurable low, medium, and high effort levels, as well as structured tool calling for agentic workflows. Its native 262K-token context can scale to a validated 1M tokens, enabling long-document processing, RAG, research, and extended coding tasks. The model was pretrained on 20T bilingual and code tokens, followed by additional mid-training and long-context training. FP8 weights reduce its memory footprint to approximately 78 GB.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    Kimi K2.7 Code

    Kimi K2.7 Code

    Coding-focused Kimi model for long-horizon agent workflows

    ...It improves end-to-end task completion across real-world programming scenarios while reducing thinking-token usage by about 30% compared with K2.6. Architecturally, it uses a 1T-parameter Mixture-of-Experts design with 32B activated parameters, 61 layers, 384 experts, a 256K-token context window, and a MoonViT vision encoder. The model supports image and video input, native INT4 quantization, interleaved thinking, and multi-step tool calling. It also forces preserve-thinking mode by default, retaining full reasoning context across multi-turn interactions to improve coding-agent consistency. K2.7 Code is recommended for use through Kimi Code CLI and can be deployed with vLLM, SGLang, or KTransformers.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    Mistral Large 3 675B Base 2512

    Mistral Large 3 675B Base 2512

    Frontier-scale 675B multimodal base model for custom AI training

    ...As the base version, it is not fine-tuned for instruction following or reasoning, making it ideal for teams planning their own domain-specific finetuning or custom training pipelines. The model is engineered for reliability, long-context comprehension, and stable performance across many enterprise, scientific, and knowledge-intensive workloads. Its architecture includes a powerful language MoE and a 2.5B-parameter vision encoder, enabling multimodal understanding out of the box. Mistral Large 3 Base supports deployment on-premises using FP8 or NVFP4 formats, enabling high-performance workflows on B200, H200, H100, or A100 hardware.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    Mistral Large 3 675B Instruct 2512 NVFP4

    Mistral Large 3 675B Instruct 2512 NVFP4

    Quantized 675B multimodal instruct model optimized for NVFP4

    ...This NVFP4 checkpoint is a post-training-activation quantized version of the original instruct model, created through a collaboration between Mistral AI, vLLM, and Red Hat using llm-compressor. It retains the same instruction-tuned behavior as the FP8 model, making it ideal for production assistants, agentic workflows, scientific tasks, and long-context enterprise systems. The model integrates a 673B-parameter MoE language backbone with a 2.5B-parameter vision encoder, enabling rich multimodal analysis across text and images. Designed for efficient deployment, it runs on a single H100 or A100 node in NVFP4 while delivering performance similar to FP8 for short- and mid-context workloads.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    Mistral Large 3 675B Instruct 2512

    Mistral Large 3 675B Instruct 2512

    Frontier-scale 675B multimodal instruct MoE model for enterprise AIMis

    ...With a 256k context window, it excels at long-document comprehension, deep retrieval workflows, and complex knowledge-intensive tasks.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    Ministral 3 3B Reasoning 2512

    Ministral 3 3B Reasoning 2512

    Compact 3B-param multimodal model for efficient on-device reasoning

    ...It supports dozens of languages, allowing it to function across global and multilingual contexts. The model retains strong system-prompt adherence, supports function-calling with structured JSON output, and offers a large 256k token context window for extended context reasoning.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17

    Statelets

    Statelets is a coordination language.

    Statelets is a coordination language addressing support and coordination of collaboration processes spanning multiple groupware tools (e.g., Wikipedia) and social networking sites (e.g., Facebook, MySpace). Statelets aim at coordination based on social and semantic network context effects, i.e., actions and intentions of related processes or people.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Alpamayo 2 Super

    Alpamayo 2 Super

    Open VLA model for autonomous driving reasoning and planning

    ...Built on the NVIDIA Cosmos platform, it enables vehicles to perceive, reason, plan, and generate driving actions using human-like decision making rather than relying solely on predefined rules. The model processes multi-camera video, navigation signals, and driving context to produce driving trajectories alongside interpretable Chain-of-Causation reasoning, improving transparency for validation and safety analysis. Alpamayo2-Super is part of the broader NVIDIA Alpamayo ecosystem, which also includes simulation frameworks, reinforcement learning infrastructure, datasets, and physical AI tools for end-to-end autonomous vehicle development. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 19
    Llama-3.2-1B

    Llama-3.2-1B

    Llama 3.2–1B: Multilingual, instruction-tuned model for mobile AI

    ...Llama 3.2-1B outperforms other open models in several benchmarks relative to its size and offers quantized versions for efficiency. It uses a refined transformer architecture with Grouped-Query Attention (GQA) and supports long context windows of up to 128k tokens.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20

    cafetiere

    Rule-based information extraction.

    UIMA-compliant text analytics using a rule language in which to express context-sensitive constraints on syntactic and semantic text elements.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Malabar is a simple, context-agnostic authentication library for Java applications. It strives to be equally usable in J2EE/Web, GUI, and other applications requiring user authentication.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Altar-1

    Altar-1

    Pruned GLM-5.3 model for self-hosted cybersecurity AI and coding

    ...Its routed experts use INT4 W4A16 AWQ quantization, while attention, shared experts, dense layers, and the output head remain BF16. The resulting weights occupy about 328 GB and are designed for production deployment on four NVIDIA H200 GPUs through vLLM, including 128K-context workloads.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Qwen3.8-27B

    Qwen3.8-27B

    Dense 27B multimodal model for coding, agents, and visual reasoning

    ...Agent capabilities emphasize autonomous planning, environment feedback, computer and browser use, and reliable completion of complex multi-step workflows. Qwen3.8-27B supports a native 262,144-token context window that can be extended to one million tokens. Thinking is enabled by default, with low, medium, and xhigh reasoning-effort settings and preserved reasoning across conversations. It also delivers substantial improvements over Qwen3.6-27B on coding and agent benchmarks while remaining deployment-friendly and compatible with Transformers, vLLM, SGLang, etc.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    Muse Glimmer

    Muse Glimmer

    Local multimodal 30B model for autonomous agents, coding, and tools

    ...Distilled from the larger Muse Spark, it combines multi-step reasoning, reliable tool use, coding, failure recovery, and image understanding in a dense 29.6B-parameter architecture with a dedicated 1.8B-parameter perception encoder. It supports more than 100 languages and a 131K+ token context window, allowing agents to maintain coherent plans across extended workflows. Muse Glimmer can interpret screenshots, charts, documents, and images alongside text, while configurable reasoning strength lets developers balance response quality and speed. Its quantized variants reduce the model below 20 GB for operation on systems with 24–32 GB of memory, and DFlash speculative decoding can substantially accelerate generation.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    Inkling

    Inkling

    Frontier multimodal MoE model for coding and AI agent workflows

    ...It contains 975B total parameters with 41B active parameters per token, balancing frontier-level capability with efficient sparse inference. The model natively processes text, images, audio, and video within a unified architecture and supports an exceptionally large 1 million token context window for long-document reasoning, repository-scale coding, and agentic execution. Trained from scratch on approximately 45 trillion multimodal tokens, Inkling introduces controllable reasoning effort, allowing users to trade off latency and reasoning depth depending on the task. It is optimized for software engineering, tool use, and large-scale autonomous workflows, with strong performance on coding and agent benchmarks.
    Downloads: 0 This Week
    Last Update:
    See Project