MSA, or Memory Sparse Attention, is a research framework for scaling language-model memory to extremely long contexts. It replaces full attention over all tokens with sparse selection of compressed latent memory states. Document-wise rotary position encoding and top-k routing keep training and inference close to linear complexity. A tiered KV-cache design stores routing keys on GPU while larger content states can remain on CPU. Its Memory Parallel engine distributes scoring and transfers only selected memory back to the accelerator. Memory Interleave alternates retrieval, context expansion, and generation to improve multi-hop reasoning across distant segments. The project reports experiments extending from 16K to 100M tokens, including inference on two A800 GPUs.

Features

  • Memory Sparse Attention architecture
  • Document-wise rotary position encoding
  • Top-k latent memory routing
  • GPU and CPU tiered KV-cache compression
  • Memory Parallel distributed inference
  • Memory Interleave for multi-hop reasoning

Project Samples

Project Activity

See All Activity >

Follow MSA: Memory Sparse Attention

MSA: Memory Sparse Attention Web Site

Other Useful Business Software
Veeam Data Platform v13.1 - Get Your Free Trial Icon
Veeam Data Platform v13.1 - Get Your Free Trial

Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
Try it Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of MSA: Memory Sparse Attention!

Additional Project Details

Programming Language

Python

Related Categories

Python Frameworks, Python Artificial Intelligence Software

Registered

1 day ago