Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Visual Causal Flow
Inference code for scalable emulation of protein equilibrium ensembles
Qwen-Image is a powerful image generation foundation model
A SOTA open-source image editing model
Use ChatGPT to summarize the arXiv papers
MOSS‑TTS Family open‑source speech and sound generation model
General-purpose image editing model that delivers high-fidelity
Multimodal Diffusion with Representation Alignment
Audio foundation model excelling in audio understanding
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Official implementation of DreamCraft3D
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Tongyi Deep Research, the Leading Open-source Deep Research Agent
OCR expert VLM powered by Hunyuan's native multimodal architecture
Open Source Speech Language Model
Video understanding codebase from FAIR for reproducing video models
Tool for exploring and debugging transformer model behaviors
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
ICLR2024 Spotlight: curation/training code, metadata, distribution
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Language modeling in a sentence representation space
Stable Diffusion WebUI Forge is a platform on top of Stable Diffusion
High-Resolution Image Synthesis with Latent Diffusion Models