Showing 23 open source projects for "cache memory simulator"

View related business solutions
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Try Free
  • Build Data Resilience - Take the Assessment Today Icon
    Build Data Resilience - Take the Assessment Today

    Can you recover when it matters most? Take this quick assessment to identify gaps and build greater recovery confidence.

    Is your recovery strategy as strong as you think? Take this quick self-assessment to check your recovery readiness and gain tailored insights. In only 2 minutes, you'll learn where you fall on the recovery readiness scale.
    Take the Assessment
  • 1
    R-KV

    R-KV

    Redundancy-aware KV Cache Compression for Reasoning Models

    R-KV is an open-source research project that focuses on improving the efficiency of large language model inference through key-value cache compression techniques. Modern transformer models rely heavily on KV caches during autoregressive decoding, which store intermediate attention states to accelerate generation. However, these caches can consume large amounts of memory, especially in reasoning-oriented models with long context windows. R-KV introduces a method for compressing the KV cache during decoding, allowing models to maintain reasoning performance while reducing memory consumption and computational overhead. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    KVCache-Factory

    KVCache-Factory

    Unified KV Cache Compression Methods for Auto-Regressive Models

    KVCache-Factory is an open-source research framework designed to explore and implement unified key-value cache compression techniques for autoregressive transformer models. In large language models, the key-value cache stores intermediate attention states that enable efficient token generation during inference, but these caches can consume large amounts of GPU memory when handling long contexts. KVCache-Factory provides a platform for implementing and evaluating multiple compression strategies that reduce memory usage while preserving model performance. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 3
    Kimi K3 in C

    Kimi K3 in C

    A 2.78-trillion-parameter Kimi K3 running inference on a single CPU

    Kimi K3 in C is a portable C99 inference engine built to run the 2.78-trillion-parameter Kimi K3 model on CPUs without BLAS, machine-learning frameworks, or GPUs. It demonstrates inference from a roughly 1.56 TB checkpoint with measured memory use as low as 8.24 GB. The runtime streams model trunk layers and routed experts from disk instead of keeping all weights resident in memory. Memory presets balance pinned layers and an expert LRU cache for laptops, desktops, workstations, and servers. Incremental generation preserves KV cache and recurrent state between tokens. ...
    Downloads: 8 This Week
    Last Update:
    See Project
  • 4
    DeepSeek-Reasonix

    DeepSeek-Reasonix

    DeepSeek-native AI coding agent for your terminal

    DeepSeek Reasonix is a DeepSeek-native AI coding agent designed for terminal-based software development. It is built around prefix-cache stability, which helps reduce token costs during long sessions and allows users to leave the agent running across extended workflows. Reasonix includes a coding mode with filesystem and shell tools, a lighter chat mode, one-shot task execution, health checks, session utilities, and project-scoped memory. It supports reviewed SEARCH/REPLACE edits, plan mode, MCP servers, web search, hooks, skills, semantic indexing, transcript replay, event logs, and cost or cache tracking. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 5
    Mooncake

    Mooncake

    Mooncake is the serving platform for Kimi

    Mooncake is an open-source infrastructure platform designed to optimize large language model serving by focusing on efficient management and transfer of model data and KV cache. The platform was originally developed as part of the serving infrastructure for the Kimi large language model system. Its architecture centers on a high-performance transfer engine that provides unified data transfer across different storage and networking technologies. This engine enables efficient movement of tensors and model data across heterogeneous environments such as GPU memory, system memory, and distributed storage systems. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 6
    Dao Code

    Dao Code

    Open-source TypeScript terminal coding agent for DeepSeek-V4

    Dao Code is an open-source terminal-native AI coding assistant built around DeepSeek V4 and a cost-conscious agent architecture. It reads code, writes code, runs commands, fixes bugs, and streams tool usage inside the terminal. The project emphasizes byte-stable prompts, prefix-cache reuse, and low-cost reflection or memory forks so longer coding sessions can remain affordable. It also includes cross-session memory that verifies saved facts against the current codebase instead of blindly trusting old context. Dao Code supports approval gates, slash commands, skills, MCP tools, hooks, custom subagents, permissions, crash recovery, shadow-git checkpoints, and autonomous long-task mode. ...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 7
    TurboFieldfare

    TurboFieldfare

    Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

    TurboFieldfare is a custom Swift and Metal runtime for running the instruction-tuned Gemma 4 26B-A4B model on Apple Silicon Macs with limited memory. Instead of loading the entire model, it keeps the shared core and KV cache in RAM while streaming only the routed experts required for each token from SSD. This approach reduces active memory use to roughly 2 GB while the installed model occupies about 14.3 GB of storage. The project includes a native Mac application, command-line tools, a streaming installer, a Swift library, and an experimental OpenAI-compatible local server. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 8
    FlashMLA

    FlashMLA

    FlashMLA: Efficient Multi-head Latent Attention Kernels

    ...It provides optimized kernels for MLA decoding, including support for variable-length sequences, helping reduce latency and increase throughput in model inference systems using that attention style. The library supports both BF16 and FP16 data types, and includes a paged KV cache implementation with a block size of 64 to efficiently manage memory during decoding. On very compute-bound settings, it can reach up to ~660 TFLOPS on H800 SXM5 hardware, while in memory-bound configurations it can push memory throughput to ~3000 GB/s. The team regularly updates it with performance improvements; for example, a 2025 update claims 5 % to 15 % gains on compute-bound workloads while maintaining API compatibility.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    Kimi Linear

    Kimi Linear

    An expressive, efficient attention architecture

    Kimi Linear is a hybrid linear attention architecture developed for efficient language modeling across short, long-context, and reinforcement learning workloads. Its core mechanism, Kimi Delta Attention, refines the gated delta rule with fine-grained controls for managing finite-state recurrent memory. The architecture combines KDA and global Multi-head Latent Attention layers at a 3:1 ratio to preserve model quality while reducing memory demands. Released Base and Instruct checkpoints contain 48 billion total parameters, activate 3 billion parameters per token, and support contexts up to one million tokens. The models were trained on 5.7 trillion tokens and can reduce KV cache requirements by up to 75 percent. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Build Securely on AWS with Proven Frameworks Icon
    Build Securely on AWS with Proven Frameworks

    Lay a foundation for success with Tested Reference Architectures developed by Fortinet’s experts. Learn more in this white paper.

    Moving to the cloud brings new challenges. How can you manage a larger attack surface while ensuring great network performance? Turn to Fortinet’s Tested Reference Architectures, blueprints for designing and securing cloud environments built by cybersecurity experts. Learn more and explore use cases in this white paper.
    Download Now
  • 10
    LMCache

    LMCache

    Supercharge Your LLM with the Fastest KV Cache Layer

    ...These capabilities aim to lower latency, cut GPU cycles, and stabilize performance for production workloads with overlapping prompts or retrieval-augmented contexts. The end result is a cache fabric for LLMs that complements engines rather than replacing them.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Colibrì

    Colibrì

    Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine

    Colibri is a compact inference engine designed to run the 744-billion-parameter GLM-5.2 mixture-of-experts model on consumer hardware. It keeps the dense portion of the quantized model in memory while streaming routed experts from a large disk-based store as they are needed. The runtime is implemented in pure C, requires no Python or BLAS during inference, and can operate without a GPU. Compressed attention caches, expert caching, optional hot tiers, and speculative decoding reduce memory...
    Downloads: 7 This Week
    Last Update:
    See Project
  • 12
    claude-obsidian

    claude-obsidian

    Claude + Obsidian knowledge companion

    claude-obsidian is an AI-powered knowledge engine that transforms an Obsidian vault into a self-organizing, continuously evolving wiki. Instead of acting as a simple chat assistant, it autonomously creates, links, and maintains structured knowledge based on user inputs and external sources. The system follows the LLM Wiki pattern, where information is stored as persistent markdown files that grow richer over time through cross-referencing and synthesis. It includes features such as...
    Downloads: 2 This Week
    Last Update:
    See Project
  • 13
    ChatLLM Web

    ChatLLM Web

    Chat with LLM like Vicuna totally in your browser with WebGPU

    ...To use this app, you need a browser that supports WebGPU, such as Chrome 113 or Chrome Canary. Chrome versions ≤ 112 are not supported. You will need a GPU with about 6.4GB of memory. If your GPU has less memory, the app will still run, but the response time will be slower. The first time you use the app, you will need to download the model. For the Vicuna-7b model that we are currently using, the download size is about 4GB. After the initial download, the model will be loaded from the browser cache for faster usage.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 14
    WASTE

    WASTE

    Run the full 2.78-trillion-parameter Kimi K3 model

    WASTE is an embeddable C inference engine for running extremely large mixture-of-experts models when the weights exceed available RAM. It keeps the shared model trunk in memory and streams only the experts selected for each token from fast NVMe storage. A bounded cache reuses recently needed experts, while lookahead routing begins disk reads before the next layer requires them. Its main target is the full 2.78-trillion-parameter Kimi K3 model, including multimodal image input. The engine has no third-party runtime dependency on its CPU inference path and exposes both a CLI and C library. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 15
    TensorRT LLM

    TensorRT LLM

    TensorRT LLM provides users with an easy-to-use Python API

    TensorRT-LLM is an open-source high-performance inference library specifically designed to optimize and accelerate large language model deployment on NVIDIA GPUs. It provides a Python-based API built on top of PyTorch that allows developers to define, customize, and deploy LLMs efficiently across a variety of hardware configurations, from single GPUs to large multi-node clusters. The library focuses on maximizing throughput and minimizing latency through advanced techniques such as...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    BigMac

    BigMac

    An open-source toolkit for BigMac-style pipeline-parallel training

    BigMac is an open-source toolkit for pipeline-parallel training of multimodal large language models. It preserves optimized language-model pipeline schedules while placing encoder and generator work around them. This design reduces activation memory without bringing back cross-module pipeline bubbles. Its scheduler creates global operator plans, while its executor runs those plans through a shared schedule abstraction. A Megatron-Core reference backend and Qwen3 and Qwen3-VL tutorials help developers connect the approach to real training workflows. The simulator lets researchers visualize schedules, compare pipeline strategies, and model timing imbalances caused by compute cost, input size, or uneven stage partitioning. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 17
    LingBot-World

    LingBot-World

    Advancing Open-source World Models

    LingBot-World is an open-source, high-fidelity world simulator designed to advance the state of world models through video generation. Built on top of Wan2.2, it enables realistic, dynamic environment simulation across diverse styles, including real-world, scientific, and stylized domains. LingBot-World supports long-term temporal consistency, maintaining coherent scenes and interactions over minute-level horizons. With real-time interactivity and sub-second latency at 16 FPS, it is...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 18
    RWKV

    RWKV

    RNN with great LLM performance

    ...The project is built around the idea that a model can be trained in a parallelizable way like a GPT-style transformer while running inference with recurrent efficiency. This gives RWKV important advantages for long-context use, including lower memory pressure and no traditional key-value cache requirement. The repository includes training code, model notes, research material, and references to current RWKV weights. Its main value is providing the foundation for experimenting with efficient large language models that combine transformer-like scalability with RNN-like runtime behavior.
    Downloads: 3 This Week
    Last Update:
    See Project
  • 19
    gpu_poor

    gpu_poor

    Calculate token/s & GPU memory requirement for any LLM

    gpu_poor is an open-source tool designed to help developers determine whether their hardware is capable of running a specific large language model and to estimate the performance they can expect from it. The project focuses on calculating GPU memory requirements and predicted inference speed for different models, hardware configurations, and quantization strategies. By analyzing factors such as model size, context length, batch size, and GPU specifications, the system estimates how much VRAM will be required and how fast tokens can be generated during inference. The tool also provides a detailed breakdown of where GPU memory is allocated, including model weights, KV cache, activations, and other runtime overhead. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 20
    The Deep Email Miner Application is a software solution for the multistaged analysis of an Email Corpus. Social network analysis and text mining techniques are connected to enable an in depth view into the underlying information. The self-executable Version 1.1 jar file will now run on Java 1.5 or higher. A Windows executable file of Version 1.1 is also provided in the Files section. Documentation can be found on the project homepage.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 21
    Origin is a evolver for the programming game Corewars. Part of this project is a modified MARS (Memory Array Redcode Simulator) written in C especialy for evolvers. With SWIG its going to be very easy to use in different languages.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Laguna XS.2

    Laguna XS.2

    Open agentic coding model optimized for local deployment

    ...The model contains 33B total parameters with only 3B activated per token, allowing it to deliver strong coding performance while remaining efficient enough to run locally on modern consumer hardware. It uses a hybrid attention architecture that combines Sliding Window Attention and global attention layers, reducing memory requirements and improving inference speed. Laguna XS.2 supports native reasoning with interleaved thinking between tool calls, enabling more capable autonomous coding agents and multi-step workflows. The model features a 262K-token context window, preserved reasoning across interactions, FP8 KV-cache optimization, and compatibility with local deployment ecosystems such as Ollama and vLLM.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    GigaChat 3 Ultra

    GigaChat 3 Ultra

    High-performance MoE model with MLA, MTP, and multilingual reasoning

    GigaChat 3 Ultra is a flagship instruct-model built on a custom Mixture-of-Experts architecture with 702B total and 36B active parameters. It leverages Multi-head Latent Attention to compress the KV cache into latent vectors, dramatically reducing memory demand and improving inference speed at scale. The model also employs Multi-Token Prediction, enabling multi-step token generation in a single pass for up to 40% faster output through speculative and parallel decoding techniques. Its training corpus incorporates ten languages, enriched with books, academic sources, code datasets, mathematical tasks, and more than 5.5 trillion tokens of high-quality synthetic data. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • Next