LTX-Video Support for ComfyUI
Text and image to video generation: CogVideoX and CogVideo
Contexts Optical Compression
gpt-oss-120b and gpt-oss-20b are two open-weight language models
Reference PyTorch implementation and models for DINOv3
Qwen3-TTS is an open-source series of TTS models
Visual Causal Flow
A Family of Open Sourced Music Foundation Models
Advanced language and coding AI model
Qwen3 is the large language model series developed by Qwen team
An Open Real-time Video-Language Interaction System
Recovering the Visual Space from Any Views
Code for running inference with the SAM 3D Body Model 3DB
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
Sharp Monocular Metric Depth in Less Than a Second
A theoretical reconstruction of the Claude Mythos architecture
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Industrial-level controllable zero-shot text-to-speech system
Unified Multimodal Understanding and Generation Models
A Powerful Native Multimodal Model for Image Generation
The official repo of Qwen chat & pretrained large language model
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Z80-μLM is a 2-bit quantized language model
Open-source image generative foundation model
Open-source multi-speaker long-form text-to-speech model