FlashKDA is an open-source library of high-performance CUDA kernels for Kimi Delta Attention, implemented on NVIDIA CUTLASS. It is intended to accelerate the forward pass used by KDA-based language models on modern NVIDIA GPUs. The package integrates with flash-linear-attention and can be selected automatically as the backend for chunk_kda during inference. It supports recurrent state input and output, variable-length batches, internal gating, query-key normalization, and beta activation. Builds can target the detected GPU architecture or multiple supported architectures for wheels and CI pipelines. The repository also provides correctness tests, benchmark material, a direct Python kernel API, and development helpers for CUDA and C++ tooling.

Features

  • High-performance Kimi Delta Attention CUDA kernels
  • NVIDIA CUTLASS-based implementation
  • Automatic flash-linear-attention backend integration
  • Stateful and variable-length sequence processing
  • Configurable GPU architecture compilation
  • Correctness tests and hardware benchmarks

Project Samples

Project Activity

See All Activity >

License

MIT License

Follow FlashKDA

FlashKDA Web Site

Other Useful Business Software
Our Free Plans just got better! | Auth0 Icon
Our Free Plans just got better! | Auth0

With up to 25k MAUs and unlimited Okta connections, our Free Plan lets you focus on what you do best—building great apps.

You asked, we delivered! Auth0 is excited to expand our Free and Paid plans to include more options so you can focus on building, deploying, and scaling applications without having to worry about your security. Auth0 now, thank yourself later.
Try free now
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of FlashKDA!

Additional Project Details

Programming Language

Python

Related Categories

Python Operating System Kernels

Registered

5 days ago