Open-source framework for intelligent speech interaction
An Open Source text-to-speech system built by inverting Whisper
Towards Human-Sounding Speech
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Speech-AI-Forge is a project developed around TTS generation model
RGBD video generation model conditioned on camera input
On-device Speech-to-Intent engine powered by deep learning
Image inpainting tool powered by SOTA AI Model
Benchmarking synthetic data generation methods
AIMET is a library that provides advanced quantization and compression
Generates original ARC-AGI-1-style tasks distribution-matched
AI cognitive-enhancement Skills based on Anthropic's J-space
Codex plugin that turns attached object images into code-only
Agent skill: make LLMs write docs in ASD-STE100
An open-source toolkit for BigMac-style pipeline-parallel training
An Open Real-time Video-Language Interaction System
Open Vision Agents by Stream. Build voice and vision agents quickly
The official Python library for the Fish Audio API
Requirement-driven evaluation harness for AI agents and LLM
AI generative media user experience highlighting use of APIs
TokenSpeed is a speed-of-light LLM inference engine
Kaggle Python docker image
A 0.1B Omni model trained from scratch
Build a modern LLM from scratch. Every line commented
Multi-source content processor for NotebookLM