Automatic Speech Recognition with Word-level Timestamps
Advanced LLM-powered brute-force tool combining AI intelligence
Open-source multi-speaker long-form text-to-speech model
A python tool that uses GPT-4, FFmpeg, and OpenCV
AI framework for automated short video creation and editing tools
Refer and Ground Anything Anywhere at Any Granularity
Long-form streaming TTS system for multi-speaker dialogue generation
Distill high-value content like books, long videos, podcasts, and more
An Open Source text-to-speech system built by inverting Whisper
A specialized Claude Code workspace for creating long-form
ClawTeam: Agent Swarm Intelligence (One Command → Full Automation)
Structured RAG: ingest, index, query
Marrying Grounding DINO with Segment Anything & Stable Diffusion
Agentic, Reasoning, and Coding (ARC) foundation models
MOSS‑TTS Family open‑source speech and sound generation model
Unleashing 10,000+ Word Generation from Long Context LLMs
Open-source, high-performance AI model with advanced reasoning
Using AI models to automatically provide commentary and edit videos
ChatGPT extension for scientific research work
Make websites accessible for AI agents
Fully Local Manus AI. No APIs, No $200 monthly bills
Claude Code blog skill suite
OCR expert VLM powered by Hunyuan's native multimodal architecture
Tools to build web AI agents that can authenticate
Pushing the Frontier of Long Audio-Visual Generation