A high-throughput and memory-efficient inference and serving engine
A lightweight vLLM implementation built from scratch
System Level Intelligent Router for Mixture-of-Models at Cloud
Welcome the Era of One-shot Long-horizon Parsing
Personal AI, On Personal Devices
GLM-5: From Vibe Coding to Agentic Engineering
A unified library of SOTA model optimization techniques
Visual Causal Flow
Private Open AI on Kubernetes
Moonshot's most powerful AI model
An expressive, efficient attention architecture
NVIDIA plugin for secure installation of OpenClaw
TokenSpeed is a speed-of-light LLM inference engine
Run a full local LLM stack with one command using Docker
Accelerate local LLM inference and finetuning
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Towards Human-Sounding Speech
From Vibe Coding to Agentic Engineering
Interface for OuteTTS models
Open source AI IDE and Cursor alternative
The free, Open Source alternative to OpenAI, Claude and others
Qwen3 is the large language model series developed by Qwen team
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs
Advanced language and coding AI model
Accurate × Fast × Comprehensive