Clone a voice in 5 seconds to generate arbitrary speech in real-time
AI PPT Track Terminator, the strongest PPT Skill ever
Miso TTS is an 8 billion, highly emotive text-to-speech model
Qwen3-omni is a natively end-to-end, omni-modal LLM
1 min voice data can also be used to train a good TTS model
The behavior guidance framework for customer-facing LLM agents
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Full git and GitHub integration with Sublime Text
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
ComfyUI wrapper nodes for HunyuanVideo
Capable of understanding text, audio, vision, video
An open-source toolkit for monitoring Language Learning Models (LLMs)
Automated translation solution for visual novels
Accurate × Fast × Comprehensive
An Open Source text-to-speech system built by inverting Whisper
A Web UI for easy subtitle using whisper model
Build Vision Agents quickly with any model or video provider
Image inpainting tool powered by SOTA AI Model
A community sourced database of game controller mappings
Parse files for optimal RAG
Open-source image generative foundation model
The official Python library for the Fish Audio API
Unlock the fullest potential of your device
Spark-TTS Inference Code
Python module for parsing semi-structured text into python tables