Models for object and human mesh reconstruction
Convert Google Gemini web into OpenAI-compatible API
LTX-Video Support for ComfyUI
Text and image to video generation: CogVideoX and CogVideo
Contexts Optical Compression
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Reference PyTorch implementation and models for DINOv3
Qwen3-TTS is an open-source series of TTS models
Visual Causal Flow
A Family of Open Sourced Music Foundation Models
Hackable and optimized Transformers building blocks
Advanced language and coding AI model
Qwen3 is the large language model series developed by Qwen team
An Open Real-time Video-Language Interaction System
Recovering the Visual Space from Any Views
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
Code for running inference with the SAM 3D Body Model 3DB
Uncommon Objects in 3D dataset
AlphaFold 3 inference pipeline
Sharp Monocular Metric Depth in Less Than a Second
A theoretical reconstruction of the Claude Mythos architecture
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Industrial-level controllable zero-shot text-to-speech system
Unified Multimodal Understanding and Generation Models
A Powerful Native Multimodal Model for Image Generation