Document Image Parsing via Heterogeneous Anchor Prompting”
Supercharge Your LLM with the Fastest KV Cache Layer
Code to accompany "A Method for Animating Children's Drawings"
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Multi-modal large language model designed for audio understanding
Real-time voice interactive digital human
Official MiniMax Model Context Protocol (MCP) server
Plug-and-play library to enable agents to call MCP and UTCP tools
Make the Chinese written by AI read like a specific person speaking
Graph-Native Infrastructure for Context and Accountable AI Systems
End-to-end protocol replay toolkit for ChatGPT Plus/Team/Pro sub
Backlog-row-first content production system for teams
AI agents running research on single-GPU nanochat training
Apple Silicon (MLX) port of Karpathy's autoresearch
Biomni: a general-purpose biomedical AI agent
A unified library of SOTA model optimization techniques
Document content and metadata extraction microservice
Machine Learning Engineering Open Book
Your Fully-Automated Personal AI Assistant
"VideoRAG: Chat with Your Videos
AI-Researcher: Autonomous Scientific Innovation
100–200× Acceleration for Video Diffusion Models
UI-TARS-desktop version that can operate on your local personal device
LLM-based agent for general purpose software engineering tasks