Easily compute clip embeddings and build a clip retrieval system
Collection of Gemma 3 variants that are trained for performance
Instant voice cloning by MIT and MyShell. Audio foundation model
An Open Source text-to-speech system built by inverting Whisper
1 min voice data can also be used to train a good TTS model
A full spaCy pipeline and models for scientific/biomedical documents
Label Studio is a multi-type data labeling and annotation tool
A modular voice assistant application for experimenting
A community-supported supercharged version of paperless
Generating Immersive, Explorable, and Interactive 3D Worlds
Framework for building real-time voice and multimodal AI agents
Underthesea - Vietnamese NLP Toolkit
AI video generator optimized for low VRAM and older GPUs use
Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model
CLIP, Predict the most relevant text snippet given an image
Python library for building agents that leverages Google Antigravity
MTEB: Massive Text Embedding Benchmark
Evaluate and monitor ML models from validation to production
An Open-Source Toolkit for General-OCR Research and Applications
Framework for building, orchestrating, and deploying AI agents
The most accurate natural language detection library for Python
Algorithms for outlier, adversarial and drift detection
Create prompt-friendly codebase digests from any Git repository URL
LLM abstractions that aren't obstructions
Parse files for optimal RAG