Experimental, AI/ML-powered and open sourced Marketing Mix Modeling
Official Repo For "Sa2VA: Marrying SAM2 with LLaVA
UI-TARS-desktop version that can operate on your local personal device
LLM-based agent for general purpose software engineering tasks
High-Fidelity and Controllable Generation of Textured 3D Assets
Multi-modal large language model designed for audio understanding
State-of-the-art (SoTA) text-to-video pre-trained model
Open-source framework for intelligent speech interaction
Large Multimodal Models for Video Understanding and Editing
A minimal yet professional single agent demo project
Real-time voice interactive digital human
An Open Source text-to-speech system built by inverting Whisper
Towards Human-Sounding Speech
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
Speech-AI-Forge is a project developed around TTS generation model
Converts text to speech in realtime
OCR expert VLM powered by Hunyuan's native multimodal architecture
RGBD video generation model conditioned on camera input
On-device Speech-to-Intent engine powered by deep learning
Benchmarking synthetic data generation methods
Making Enterprise Data Intelligent and Responsive for AI
AIMET is a library that provides advanced quantization and compression
Powering Amazon custom machine learning chips
Miso TTS is an 8 billion, highly emotive text-to-speech model
An advanced paper search agent powered by large language models