A robust, efficient, low-latency speech-to-text library
State-of-the-art (SoTA) text-to-video pre-trained model
Open Source Document Management System for Digital Archives
The official Python SDK for the ElevenLabs API
A fast TTS architecture with conditional flow matching
The behavior guidance framework for customer-facing LLM agents
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
The simplest, fastest repository for training/finetuning models
Free, high-quality text-to-speech API endpoint to replace OpenAI
ComfyUI wrapper nodes for HunyuanVideo
Capable of understanding text, audio, vision, video
An open-source toolkit for monitoring Language Learning Models (LLMs)
Automated translation solution for visual novels
Accurate × Fast × Comprehensive
Open source machine learning framework to automate text conversations
Agent harness to make your slop code well-engineered and beautiful
Spark-TTS Inference Code
GLM-4-Voice | End-to-End Chinese-English Conversational Model
A Unified Framework for Text-to-3D and Image-to-3D Generation
Easy-to-use and powerful NLP library with Awesome model zoo
Designed for text embedding and ranking tasks
Generating Immersive, Explorable, and Interactive 3D Worlds
Foundational video generation model with 13.6B parameters
Persian NLP Toolkit
Easily compute clip embeddings and build a clip retrieval system