Reference PyTorch implementation and models for DINOv3
A Unified Framework for Text-to-3D and Image-to-3D Generation
Open-source deep-learning framework
Open Source Speech Language Model
Python inference and LoRA trainer package for the LTX-2 audio–video
FAIR Sequence Modeling Toolkit 2
1B text generation model based on the HRM architecture
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Official repository for LTX-Video
Project Lyra: Open Generative 3D World Models
Text and image to video generation: CogVideoX and CogVideo
High-resolution models for human tasks
Block Diffusion for Ultra-Fast Speculative Decoding
LTX-Video Support for ComfyUI
A Powerful Native Multimodal Model for Image Generation
Generating Immersive, Explorable, and Interactive 3D Worlds
A Systematic Framework for Interactive World Modeling
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Achieving 3+ generation speedup on reasoning tasks
Open-Source Financial Large Language Models
Video understanding codebase from FAIR for reproducing video models
An Efficient Agentic Model for Computer Use
Unified Multimodal Understanding and Generation Models
PyTorch code and models for the DINOv2 self-supervised learning
CogView4, CogView3-Plus and CogView3(ECCV 2024)