State-of-the-art TTS model under 25MB
Open-source, high-performance AI model with advanced reasoning
Powerful AI language model (MoE) optimized for efficiency/performance
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Awesome multilingual OCR toolkits based on PaddlePaddle
Native and Compact Structured Latents for 3D Generation
Visual Causal Flow
An Open Real-time Video-Language Interaction System
Open-source multi-speaker long-form text-to-speech model
Industrial-level controllable zero-shot text-to-speech system
From Images to High-Fidelity 3D Assets
AlphaFold 3 inference pipeline
Video understanding codebase from FAIR for reproducing video models
Long-form streaming TTS system for multi-speaker dialogue generation
Video Object and Interaction Deletion
A theoretical reconstruction of the Claude Mythos architecture
Robust Speech Recognition Across Languages, Dialects
Contexts Optical Compression
General-purpose image editing model that delivers high-fidelity
Inference script for Oasis 500M
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Audio foundation model excelling in audio understanding
Open Source Speech Language Model
code for Mesh R-CNN, ICCV 2019