Running a big model on a small laptop
MiMo-V2-Flash: Efficient Reasoning, Coding, and Agentic Foundation
Fast, Sharp & Reliable Agentic Intelligence
Port of OpenAI's Whisper model in C/C++
DeepSeek 4 Flash local inference engine for Metal
Running a 28.9M parameter LLM on an $8 microcontroller
Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI
The official Python SDK for the ElevenLabs API
20+ high-performance LLMs with recipes to pretrain, finetune at scale
Convert Google Gemini web into OpenAI-compatible API
Unified KV Cache Compression Methods for Auto-Regressive Models
AI video skill for Claude Code & Codex
Open-source large language model family from Tencent Hunyuan
OpenBot leverages smartphones as brains for low-cost robots
Deploy your private Gemini application for free with one click
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
CodeGeeX2: A More Powerful Multilingual Code Generation Model
An opinionated CLI to transcribe Audio files w/ Whisper on-device
Open-source AI video pipeline, fully automated with MCP
Microservices-based flash sale system for high-concurrency testing
Flash enables you to easily configure and run complex AI recipes
Deep learning PyTorch library for time series forecasting
A JavaScript HTML screenshot renderer