From Images to High-Fidelity 3D Assets
Robust Speech Recognition Across Languages, Dialects
Phi-3.5 for Mac: Locally-run Vision and Language Models
High-Resolution Image Synthesis with Latent Diffusion Models
Convert Google Gemini web into OpenAI-compatible API
26m function call model that runs on incredibly small devices
State-of-the-art TTS model under 25MB
Open-source image generative foundation model
Tencent Hunyuan Multimodal diffusion transformer (MM-DiT) model
Personalize Any Characters with a Scalable Diffusion Transformer
Qwen-Image is a powerful image generation foundation model
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference
A Powerful Native Multimodal Model for Image Generation
An Efficient Agentic Model for Computer Use
Code for running inference and finetuning with SAM 3 model
Multimodal-Driven Architecture for Customized Video Generation
A Pragmatic VLA Foundation Model
Collection of Gemma 3 variants that are trained for performance
Qwen3-TTS is an open-source series of TTS models
Open image model at the forefront of design
Lets make video diffusion practical
Powerful AI language model (MoE) optimized for efficiency/performance
HY-Motion model for 3D character animation generation
Pokee Deep Research Model Open Source Repo
Visual Causal Flow