MobileLLM Optimizing Sub-billion Parameter Language Models
Redundancy-aware KV Cache Compression for Reasoning Models
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Learn AI and LLMs from scratch using free resources
Calculate token/s & GPU memory requirement for any LLM
Frontier multimodal MoE model for coding and AI agent workflows