FlashKDA is an open-source library of high-performance CUDA kernels for Kimi Delta Attention, implemented on NVIDIA CUTLASS. It is intended to accelerate the forward pass used by KDA-based language models on modern NVIDIA GPUs. The package integrates with flash-linear-attention and can be selected automatically as the backend for chunk_kda during inference. It supports recurrent state input and output, variable-length batches, internal gating, query-key normalization, and beta activation. Builds can target the detected GPU architecture or multiple supported architectures for wheels and CI pipelines. The repository also provides correctness tests, benchmark material, a direct Python kernel API, and development helpers for CUDA and C++ tooling.

Features

  • High-performance Kimi Delta Attention CUDA kernels
  • NVIDIA CUTLASS-based implementation
  • Automatic flash-linear-attention backend integration
  • Stateful and variable-length sequence processing
  • Configurable GPU architecture compilation
  • Correctness tests and hardware benchmarks

Project Samples

Project Activity

See All Activity >

License

MIT License

Follow FlashKDA

FlashKDA Web Site

Other Useful Business Software
Go from Code to Production URL in Seconds Icon
Go from Code to Production URL in Seconds

Cloud Run deploys apps in any language instantly. Scales to zero. Pay only when code runs.

Skip the Kubernetes configs. Cloud Run handles HTTPS, scaling, and infrastructure automatically. Two million requests free per month.
Try it free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of FlashKDA!

Additional Project Details

Programming Language

Python

Related Categories

Python Operating System Kernels

Registered

6 days ago