Official repository for LTX-Video
Recovering the Visual Space from Any Views
LTX-Video Support for ComfyUI
From Vibe Coding to Agentic Engineering
Video Object and Interaction Deletion
Contexts Optical Compression
Bidirectional token-classification model for identifiable info
Visual Causal Flow
Implementation of the Surya Foundation Model for Heliophysics
Open-source multi-speaker long-form text-to-speech model
A theoretical reconstruction of the Claude Mythos architecture
Ultra-Efficient LLMs on End Device
Sharp Monocular Metric Depth in Less Than a Second
Video understanding codebase from FAIR for reproducing video models
Audio foundation model excelling in audio understanding
Official implementation of DreamCraft3D
Your clothes, extracted and organized with gpt-image
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Open Source Speech Language Model
Multimodal model achieving SOTA performance
Diffusion Transformer with Fine-Grained Chinese Understanding
Large Multimodal Models for Video Understanding and Editing
Large-language-model & vision-language-model based on Linear Attention
Encoder of greater-than-word length text trained on a variety of data
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation