Implement a concise and clear Deep Search Agent from 0
The best ChatGPT that $100 can buy
4M: Massively Multimodal Masked Modeling
ICLR2024 Spotlight: curation/training code, metadata, distribution
[CVPR 2025 Best Paper Award] VGGT
A Customizable Image-to-Video Model based on HunyuanVideo
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
RGBD video generation model conditioned on camera input
Unifying 3D Mesh Generation with Language Models
AI agents autonomously run and improve ML experiments overnight
A personal context-agent that learns how you work
Controllable & emotion-expressive zero-shot TTS
Controllable and fast Text-to-Speech for over 7000 languages
PyTorch code and models for VJEPA2 self-supervised learning from video
Educational framework exploring multi-agent orchestration
This repo contains the code for 1D tokenizer and generator
Flexible Photo Recrafting While Preserving Your Identity
Taming Stable Diffusion for Lip Sync
Multi-Agent daTa geneRation Infra and eXperimentation framework
Build cross-modal and multimodal applications on the cloud
A set of Docker images for training and serving models in TensorFlow
GUI Exploration Lab. One of the best GUI agent solutions
Large-language-model & vision-language-model based on Linear Attention
Open source demo platform where you can easily showcase your AI models