Self-hosted game stream host for Moonlight
FlashInfer: Kernel Library for LLM Serving
A robust, efficient, low-latency speech-to-text library
MII makes low-latency and high-throughput inference possible
Faster Whisper transcription with CTranslate2
Compress tool outputs, logs, files, and RAG chunks
Personal AI, On Personal Devices
Ffree local self hosted video compressor webui
Open-Source Low-Latency Accelerated Linux WebRTC HTML5 Remote Desktop
Optimizing inference proxy for LLMs
MOSS-TTS-Nano is an open-source multilingual tiny speech generation
Machine learning on FPGAs using HLS
RF-DETR is a real-time object detection and segmentation
Build Vision Agents quickly with any model or video provider
Fast backend for long-term AI user memory via structured profiles
Towards Human-Sounding Speech
AI memory OS for LLM and Agent systems
Blazing-fast vector DB with similarity search and metadata filtering
Advancing Open-source World Models
NVR with realtime local object detection for IP cameras
Fast multimodal LLM for real-time voice interaction and AI apps
Low-latency AI inference engine optimized for mobile devices
Cache-Augmented Generation: A Simple, Efficient Alternative to RAG
LightLLM is a Python-based LLM (Large Language Model) inference
Open Vision Agents by Stream. Build voice and vision agents quickly