This repo contains the code for 1D tokenizer and generator
Qwen3-omni is a natively end-to-end, omni-modal LLM
Code for running inference and finetuning with SAM 3 model
A 0.1B Omni model trained from scratch
Focus on prompting and generating
InvokeAI is a leading creative engine for Stable Diffusion models
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
A Multi-Modal World Model for Reconstructing, Generating, Simulation
High-Resolution Image Synthesis with Latent Diffusion Models
AI generative media user experience highlighting use of APIs
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Free, high-quality text-to-speech API endpoint to replace OpenAI
Accurate × Fast × Comprehensive
State-of-the-art diffusion models for image and audio generation
Dealing with all unstructured data, such as reverse image search
Parse files for optimal RAG
Multilingual sentence & image embeddings with BERT
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
ComfyUI wrapper nodes for WanVideo and related models
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System
Unified Multimodal Understanding and Generation Models
Fast-stable-diffusion + DreamBooth
"Big Model" trains a visual multimodal VLM with 26M parameters
Implementation of "MobileCLIP" CVPR 2024
Multimodal embedding and reranking models built on Qwen3-VL