FlashMLA: Efficient Multi-head Latent Attention Kernels
Powerful AI language model (MoE) optimized for efficiency/performance
A theoretical reconstruction of the Claude Mythos architecture
An expressive, efficient attention architecture
Build your own Cowork, AI Scientist and other SoTA Agents
Desktop research workspace for PDFs, notes, citations, bibliographies.
Open-source KSI mappings + OSCAL examples for FedRAMP 20x.
Lightweight MoE model for local reasoning, coding, and AI agents
High-performance MoE model with MLA, MTP, and multilingual reasoning