MSA, or Memory Sparse Attention, is a research framework for scaling language-model memory to extremely long contexts. It replaces full attention over all tokens with sparse selection of compressed latent memory states. Document-wise rotary position encoding and top-k routing keep training and inference close to linear complexity. A tiered KV-cache design stores routing keys on GPU while larger content states can remain on CPU. Its Memory Parallel engine distributes scoring and transfers only selected memory back to the accelerator. Memory Interleave alternates retrieval, context expansion, and generation to improve multi-hop reasoning across distant segments. The project reports experiments extending from 16K to 100M tokens, including inference on two A800 GPUs.

Features

  • Memory Sparse Attention architecture
  • Document-wise rotary position encoding
  • Top-k latent memory routing
  • GPU and CPU tiered KV-cache compression
  • Memory Parallel distributed inference
  • Memory Interleave for multi-hop reasoning

Project Samples

Project Activity

See All Activity >

Follow MSA: Memory Sparse Attention

MSA: Memory Sparse Attention Web Site

Other Useful Business Software
Earn up to 16% annual interest with Nexo. Icon
Earn up to 16% annual interest with Nexo.

Access competitive interest rates on your digital assets.

Generate interest, borrow against your crypto, and trade a range of cryptocurrencies — all in one platform. Geographic restrictions, eligibility, and terms apply.
Get started with Nexo.
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of MSA: Memory Sparse Attention!

Additional Project Details

Programming Language

Python

Related Categories

Python Frameworks, Python Artificial Intelligence Software

Registered

1 day ago