Multimodal-Driven Architecture for Customized Video Generation
From Images to High-Fidelity 3D Assets
Multimodal Diffusion with Representation Alignment
ICLR2024 Spotlight: curation/training code, metadata, distribution
Hackable and optimized Transformers building blocks
Official implementation of DreamCraft3D
An experimental version of DeepSeek model
Open-Source Financial Large Language Models
An Efficient Agentic Model for Computer Use
The official PyTorch implementation of Google's Gemma models
code for Mesh R-CNN, ICCV 2019
A Production-ready Reinforcement Learning AI Agent Library
Diffusion Transformer with Fine-Grained Chinese Understanding
Pokee Deep Research Model Open Source Repo
The ChatGPT Retrieval Plugin lets you easily find personal documents
Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference implementation of the Transformer architecture optimized
Learning to Act by Watching Unlabeled Online Videos