Translate the video from one language to another and embed dubbing
Qwen3-omni is a natively end-to-end, omni-modal LLM
Towards Human-Sounding Speech
Multimodal-Driven Architecture for Customized Video Generation
Python module for parsing semi-structured text into python tables
Miso TTS is an 8 billion, highly emotive text-to-speech model
A Web UI for easy subtitle using whisper model
Video-based AI memory library. Store millions of text chunks in MP4
Collection of Gemma 3 variants that are trained for performance
Unifying 3D Mesh Generation with Language Models
AI PPT Track Terminator, the strongest PPT Skill ever
The official Python library for the Fish Audio API
Framework for building realtime multimodal voice AI agents apps
Open source healthcare AI
Pycorrector is a toolkit for text error correction
Spark-TTS Inference Code
Free, high-quality text-to-speech API endpoint to replace OpenAI
Generate blog articles from video or audio
On-device TTS model by Neuphonic
A high-quality PDF to Markdown tool based on large language model
A high-quality rapid TTS voice cloning model
An easy-to-use backup tool for GNU Linux using rsync in the back
OCR model for complex documents with layout-aware structured outputs
Image inpainting tool powered by SOTA AI Model
A python parametric CAD scripting framework based on OCCT