FlashInfer: Kernel Library for LLM Serving
A robust, efficient, low-latency speech-to-text library
MII makes low-latency and high-throughput inference possible
Faster Whisper transcription with CTranslate2
Compress tool outputs, logs, files, and RAG chunks
Personal AI, On Personal Devices
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Optimizing inference proxy for LLMs
Machine learning on FPGAs using HLS
RF-DETR is a real-time object detection and segmentation
Build Vision Agents quickly with any model or video provider
Fast backend for long-term AI user memory via structured profiles
AI memory OS for LLM and Agent systems
Towards Human-Sounding Speech
Advancing Open-source World Models
NVR with realtime local object detection for IP cameras
Fast multimodal LLM for real-time voice interaction and AI apps
Low-latency AI inference engine optimized for mobile devices
Cache-Augmented Generation: A Simple, Efficient Alternative to RAG
LightLLM is a Python-based LLM (Large Language Model) inference
Open Vision Agents by Stream. Build voice and vision agents quickly
Parallax is a distributed model serving framework
Implementation of "MobileCLIP" CVPR 2024
The official Python SDK for the ElevenLabs API
Deep learning optimization library: makes distributed training easy