Video Object and Interaction Deletion
Official implementation of Watermark Anything with Localized Messages
High-resolution models for human tasks
Video understanding codebase from FAIR for reproducing video models
Multimodal-Driven Architecture for Customized Video Generation
Qwen-Image is a powerful image generation foundation model
Robust Speech Recognition Across Languages, Dialects
Tiny vision language model
Programmatic access to the AlphaGenome model
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
scikit-learn compatible tabular foundation model
Model export recipes, Python primitives, and Swift runtime utilities
Long-form streaming TTS system for multi-speaker dialogue generation
General-purpose image editing model that delivers high-fidelity
PyTorch code and models for the DINOv2 self-supervised learning
A Customizable Image-to-Video Model based on HunyuanVideo
Open-source large language model family from Tencent Hunyuan
Personalize Any Characters with a Scalable Diffusion Transformer
C++ implementation of ChatGLM-6B & ChatGLM2-6B & ChatGLM3 & GLM4(V)
Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
GLM-4-Voice | End-to-End Chinese-English Conversational Model
Phi-3.5 for Mac: Locally-run Vision and Language Models
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
State-of-the-art (SoTA) text-to-video pre-trained model
1B text generation model based on the HRM architecture