Showing 51 open source projects for "cache"

View related business solutions
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • 1
    DeepSeek-Reasonix

    DeepSeek-Reasonix

    DeepSeek-native AI coding agent for your terminal

    ...The project is especially useful for developers who want an open, terminal-first coding agent optimized for DeepSeek’s cache mechanics. It also includes a prerelease desktop client for users who prefer a GUI over the same agent loop.
    Downloads: 99 This Week
    Last Update:
    See Project
  • 2
    LMCache

    LMCache

    Supercharge Your LLM with the Fastest KV Cache Layer

    ...These capabilities aim to lower latency, cut GPU cycles, and stabilize performance for production workloads with overlapping prompts or retrieval-augmented contexts. The end result is a cache fabric for LLMs that complements engines rather than replacing them.
    Downloads: 23 This Week
    Last Update:
    See Project
  • 3
    R-KV

    R-KV

    Redundancy-aware KV Cache Compression for Reasoning Models

    R-KV is an open-source research project that focuses on improving the efficiency of large language model inference through key-value cache compression techniques. Modern transformer models rely heavily on KV caches during autoregressive decoding, which store intermediate attention states to accelerate generation. However, these caches can consume large amounts of memory, especially in reasoning-oriented models with long context windows. R-KV introduces a method for compressing the KV cache during decoding, allowing models to maintain reasoning performance while reducing memory consumption and computational overhead. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    KVCache-Factory

    KVCache-Factory

    Unified KV Cache Compression Methods for Auto-Regressive Models

    KVCache-Factory is an open-source research framework designed to explore and implement unified key-value cache compression techniques for autoregressive transformer models. In large language models, the key-value cache stores intermediate attention states that enable efficient token generation during inference, but these caches can consume large amounts of GPU memory when handling long contexts. KVCache-Factory provides a platform for implementing and evaluating multiple compression strategies that reduce memory usage while preserving model performance. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    CAG

    CAG

    Cache-Augmented Generation: A Simple, Efficient Alternative to RAG

    CAG, or Cache-Augmented Generation, is an experimental framework that explores an alternative architecture for integrating external knowledge into large language model responses. Traditional retrieval-augmented generation systems rely on real-time retrieval of documents from databases or vector stores during inference. CAG proposes a different approach by preloading relevant knowledge into the model’s context window and precomputing the model’s key-value cache before queries are processed. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    Kimi K3 in C

    Kimi K3 in C

    A 2.78-trillion-parameter Kimi K3 running inference on a single CPU

    ...The runtime streams model trunk layers and routed experts from disk instead of keeping all weights resident in memory. Memory presets balance pinned layers and an expert LRU cache for laptops, desktops, workstations, and servers. Incremental generation preserves KV cache and recurrent state between tokens. The repository also includes tokenization, safetensors loading, diagnostics, benchmarks, trace replay, and comparisons against a PyTorch reference.
    Downloads: 12 This Week
    Last Update:
    See Project
  • 7
    FastDeploy

    FastDeploy

    High-performance Inference and Deployment Toolkit for LLMs and VLMs

    ...The platform enables developers to deploy trained models quickly using optimized inference pipelines that support GPUs, specialized AI accelerators, and other hardware architectures. FastDeploy includes advanced acceleration technologies such as speculative decoding, multi-token prediction, and efficient KV cache management to improve throughput and latency during inference. It also offers compatibility with OpenAI-style APIs and vLLM-like interfaces, allowing developers to integrate deployed models easily into existing applications and services.
    Downloads: 15 This Week
    Last Update:
    See Project
  • 8
    Mooncake

    Mooncake

    Mooncake is the serving platform for Kimi

    ...Mooncake also introduces distributed key-value cache storage that allows inference systems to reuse previously computed attention states, significantly improving throughput in large-scale deployments. The system supports advanced networking technologies such as RDMA and NVMe over Fabric, enabling high-speed communication across clusters.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    TurboFieldfare

    TurboFieldfare

    Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

    ...Quantized weights, custom Metal kernels, chunked prefill, and a bounded expert cache improve efficiency. The installer downloads and repacks model ranges without staging a second full checkpoint. Its current scope is text-only Gemma inference on macOS 26 and Apple Silicon hardware.
    Downloads: 4 This Week
    Last Update:
    See Project
  • 99.99% Uptime for MySQL and PostgreSQL Databases Icon
    99.99% Uptime for MySQL and PostgreSQL Databases

    Sub-second maintenance. 2x read/write performance. Built-in vector search for AI apps.

    Cloud SQL Enterprise Plus delivers near-zero downtime with 35 days of point-in-time recovery. Supports MySQL, PostgreSQL, and SQL Server.
    Start Free
  • 10
    Codex-X

    Codex-X

    Codex Switch & Instruct desktop manager

    ...It replaces repeated manual file editing with a visual interface for prompts, providers, sessions, skills, MCP servers, and TOML settings. Users can categorize, import, edit, enable, disable, cache, and synchronize Markdown instruction templates. Provider tools store multiple API configurations, test connections, retrieve models, and switch between official and third-party services. Session management can search, group, inspect, synchronize, and permanently delete local Codex histories. The application also displays authentication and configuration files while creating backups before important changes. ...
    Downloads: 39 This Week
    Last Update:
    See Project
  • 11
    MSA: Memory Sparse Attention

    MSA: Memory Sparse Attention

    Trainable latent-memory framework for 100M-token contexts

    ...It replaces full attention over all tokens with sparse selection of compressed latent memory states. Document-wise rotary position encoding and top-k routing keep training and inference close to linear complexity. A tiered KV-cache design stores routing keys on GPU while larger content states can remain on CPU. Its Memory Parallel engine distributes scoring and transfers only selected memory back to the accelerator. Memory Interleave alternates retrieval, context expansion, and generation to improve multi-hop reasoning across distant segments. The project reports experiments extending from 16K to 100M tokens, including inference on two A800 GPUs.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    JDA

    JDA

    Java wrapper for the popular chat & VOIP service

    ...After setting the token and other options via setters, the JDA Object is then created by calling the build() method. When build() returns, JDA might not have finished starting up. However, you can use await ready() on the JDA object to ensure that the entire cache is loaded before proceeding.
    Downloads: 4 This Week
    Last Update:
    See Project
  • 13
    Kimi Linear

    Kimi Linear

    An expressive, efficient attention architecture

    ...Released Base and Instruct checkpoints contain 48 billion total parameters, activate 3 billion parameters per token, and support contexts up to one million tokens. The models were trained on 5.7 trillion tokens and can reduce KV cache requirements by up to 75 percent. Kimi Linear supports inference through Hugging Face Transformers and deployment as an OpenAI-compatible API with vLLM.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    ds4.c

    ds4.c

    DeepSeek 4 Flash local inference engine for Metal

    ...Unlike general-purpose inference runtimes, the project is intentionally optimized for a specific model family, enabling highly efficient execution and simplified architecture. The engine includes DS4-specific model loading, KV cache management, prompt rendering, and OpenAI-compatible server APIs for local deployment workflows. Built as a native low-level implementation, it focuses on performance, reduced abstraction overhead, and direct integration with Apple GPU acceleration through Metal compute graphs. The project also supports streaming inference behavior and local API serving for integration with external tools and AI applications. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    oMLX

    oMLX

    LLM inference server with continuous batching & SSD caching

    ...It serves text models, vision-language models, OCR models, embeddings, and rerankers through OpenAI- and Anthropic-compatible APIs. Continuous batching allows concurrent requests, while a tiered KV cache keeps active data in RAM and moves colder blocks to SSD for reuse. Models can be pinned, unloaded manually, or evicted automatically when memory runs low. A web dashboard provides monitoring, chat, downloads, benchmarks, integrations, and per-model settings. It also supports tool calling, structured output, MCP integration, and experimental multi-Mac inference.
    Downloads: 15 This Week
    Last Update:
    See Project
  • 16
    Dao Code

    Dao Code

    Open-source TypeScript terminal coding agent for DeepSeek-V4

    Dao Code is an open-source terminal-native AI coding assistant built around DeepSeek V4 and a cost-conscious agent architecture. It reads code, writes code, runs commands, fixes bugs, and streams tool usage inside the terminal. The project emphasizes byte-stable prompts, prefix-cache reuse, and low-cost reflection or memory forks so longer coding sessions can remain affordable. It also includes cross-session memory that verifies saved facts against the current codebase instead of blindly trusting old context. Dao Code supports approval gates, slash commands, skills, MCP tools, hooks, custom subagents, permissions, crash recovery, shadow-git checkpoints, and autonomous long-task mode. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    UCCL

    UCCL

    UCCL is an efficient communication library for GPUs

    ...UCCL is designed to work with heterogeneous hardware environments, allowing GPUs from different vendors and network interfaces to communicate efficiently without vendor lock-in. The system also supports specialized workloads such as reinforcement learning weight transfers, key-value cache sharing, and expert parallelism for mixture-of-experts models. Its architecture emphasizes flexibility and extensibility so that developers can implement custom communication protocols tailored to specific machine learning workloads.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    FlashMLA

    FlashMLA

    FlashMLA: Efficient Multi-head Latent Attention Kernels

    ...It provides optimized kernels for MLA decoding, including support for variable-length sequences, helping reduce latency and increase throughput in model inference systems using that attention style. The library supports both BF16 and FP16 data types, and includes a paged KV cache implementation with a block size of 64 to efficiently manage memory during decoding. On very compute-bound settings, it can reach up to ~660 TFLOPS on H800 SXM5 hardware, while in memory-bound configurations it can push memory throughput to ~3000 GB/s. The team regularly updates it with performance improvements; for example, a 2025 update claims 5 % to 15 % gains on compute-bound workloads while maintaining API compatibility.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Grok CLI

    Grok CLI

    An open-source AI agent that brings the power of Grok

    Grok CLI is a command-line interface built around the Grok AI model that brings programmatic and conversational AI capabilities directly to developer terminals. It lets you run Grok queries from your shell, scripting environment, or automation workflows without switching to a browser, enabling utility in scripting, quick data exploration, code generation, and assistant-guided tasks directly where you write code. The CLI supports streaming responses, so outputs appear in real time as the Grok...
    Downloads: 26 This Week
    Last Update:
    See Project
  • 20
    ContextForge MCP Gateway

    ContextForge MCP Gateway

    A Model Context Protocol (MCP) Gateway & Registry

    MCP Context Forge is a feature-rich gateway and registry that federates Model Context Protocol (MCP) servers and traditional REST services behind a single, governed endpoint. It exposes an MCP-compliant interface to clients while handling discovery, authentication, rate limiting, retries, and observability on the server side. The gateway scales horizontally, supports multi-cluster deployments on Kubernetes, and uses Redis for federation and caching across instances. Operators can define...
    Downloads: 8 This Week
    Last Update:
    See Project
  • 21
    mac code

    mac code

    Claude Code, but it runs on your Mac for free

    mac code is a local AI coding agent designed to run large language models directly on Apple Silicon machines without relying on cloud services, effectively transforming a Mac into a self-contained AI development environment. The project focuses on enabling models that traditionally exceed available RAM to run efficiently by streaming model weights from SSD storage, thereby overcoming hardware limitations through innovative memory management techniques. It operates as a CLI-based assistant...
    Downloads: 4 This Week
    Last Update:
    See Project
  • 22
    claude-obsidian

    claude-obsidian

    Claude + Obsidian knowledge companion

    claude-obsidian is an AI-powered knowledge engine that transforms an Obsidian vault into a self-organizing, continuously evolving wiki. Instead of acting as a simple chat assistant, it autonomously creates, links, and maintains structured knowledge based on user inputs and external sources. The system follows the LLM Wiki pattern, where information is stored as persistent markdown files that grow richer over time through cross-referencing and synthesis. It includes features such as...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 23
    ChatLLM Web

    ChatLLM Web

    Chat with LLM like Vicuna totally in your browser with WebGPU

    ...The first time you use the app, you will need to download the model. For the Vicuna-7b model that we are currently using, the download size is about 4GB. After the initial download, the model will be loaded from the browser cache for faster usage.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 24
    tiny-llm

    tiny-llm

    A course of learning LLM inference serving on Apple Silicon

    tiny-llm is an educational open-source project designed to teach system engineers how large language model inference and serving systems work by building them from scratch. The project is structured as a guided course that walks developers through the process of implementing the core components required to run a modern language model, including attention mechanisms, token generation, and optimization techniques. Rather than relying on high-level machine learning frameworks, the codebase uses...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 25
    Agent Skills Tech Leads Club

    Agent Skills Tech Leads Club

    The secure, validated skill registry for professional AI coding agents

    ...The project emphasizes security through open-source content, static analysis, content hashing, lockfiles, path isolation, and audit logging. Its CLI can browse, install, update, remove, cache, and inspect skills at local or global scope. An interactive wizard helps users choose skills, target agents, and installation methods. A companion MCP server exposes the catalog through search-first, on-demand retrieval. The registry covers development, cloud, browser automation, design, security review, and other engineering tasks.
    Downloads: 1 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • 3
  • Next