Fast-stable-diffusion + DreamBooth
A trainable PyTorch reproduction of AlphaFold 3
Generates original ARC-AGI-1-style tasks distribution-matched
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Genome modeling and design across all domains of life
Unified Multimodal Understanding and Generation Models
Use ChatGPT to summarize the arXiv papers
Open image model at the forefront of design
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Inference script for Oasis 500M
Fast and Universal 3D reconstruction model for versatile tasks
Open Source Speech Language Model
Open-source industrial-grade ASR models
Implementation of "MobileCLIP" CVPR 2024
High-resolution models for human tasks
Ling is a MoE LLM provided and open-sourced by InclusionAI
MOSS‑TTS Family open‑source speech and sound generation model
High-Fidelity and Controllable Generation of Textured 3D Assets
OCR expert VLM powered by Hunyuan's native multimodal architecture
Advancing Open-source World Models
A Systematic Framework for Interactive World Modeling
Reproduction of Poetiq's record-breaking submission to the ARC-AGI-1
DeepMind model for tracking arbitrary points across videos & robotics