FlashKDA is an open-source library of high-performance CUDA kernels for Kimi Delta Attention, implemented on NVIDIA CUTLASS. It is intended to accelerate the forward pass used by KDA-based language models on modern NVIDIA GPUs. The package integrates with flash-linear-attention and can be selected automatically as the backend for chunk_kda during inference. It supports recurrent state input and output, variable-length batches, internal gating, query-key normalization, and beta activation. Builds can target the detected GPU architecture or multiple supported architectures for wheels and CI pipelines. The repository also provides correctness tests, benchmark material, a direct Python kernel API, and development helpers for CUDA and C++ tooling.

Features

  • High-performance Kimi Delta Attention CUDA kernels
  • NVIDIA CUTLASS-based implementation
  • Automatic flash-linear-attention backend integration
  • Stateful and variable-length sequence processing
  • Configurable GPU architecture compilation
  • Correctness tests and hardware benchmarks

Project Samples

Project Activity

See All Activity >

License

MIT License

Follow FlashKDA

FlashKDA Web Site

Other Useful Business Software
$300 Free Credits to Build on Google Cloud Icon
$300 Free Credits to Build on Google Cloud

New customers can spin up VMs, build with AI, and query data at no cost.

Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of FlashKDA!

Additional Project Details

Programming Language

Python

Related Categories

Python Operating System Kernels

Registered

2026-07-29