LTX-Video Support for ComfyUI
Models for object and human mesh reconstruction
Text and image to video generation: CogVideoX and CogVideo
Convert Google Gemini web into OpenAI-compatible API
Contexts Optical Compression
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Reference PyTorch implementation and models for DINOv3
Qwen3-TTS is an open-source series of TTS models
Visual Causal Flow
A Family of Open Sourced Music Foundation Models
Advanced language and coding AI model
Hackable and optimized Transformers building blocks
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
Qwen3 is the large language model series developed by Qwen team
An Open Real-time Video-Language Interaction System
Recovering the Visual Space from Any Views
Sharp Monocular Metric Depth in Less Than a Second
Code for running inference with the SAM 3D Body Model 3DB
AlphaFold 3 inference pipeline
Uncommon Objects in 3D dataset
A theoretical reconstruction of the Claude Mythos architecture
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Industrial-level controllable zero-shot text-to-speech system
Unified Multimodal Understanding and Generation Models
A Powerful Native Multimodal Model for Image Generation