Qwen-Image is a powerful image generation foundation model
Guiding Instruction-based Image Editing via Multimodal Large Language
Capable of understanding text, audio, vision, video
CogView4, CogView3-Plus and CogView3(ECCV 2024)
AI-powered code assistant for Vim. OpenAI and ChatGPT plugin for Vim
Tensor search for humans
Qwen3-omni is a natively end-to-end, omni-modal LLM
Multilingual sentence & image embeddings with BERT
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System
Phi-3.5 for Mac: Locally-run Vision and Language Models
Large-language-model & vision-language-model based on Linear Attention
Open source libraries and APIs to build custom preprocessing pipelines
Gemma open-weight LLM library, from Google DeepMind
Chinese and English multimodal conversational language model
LISA: Reasoning Segmentation via Large Language Model
Skywork-R1V is an advanced multimodal AI model series
Refer and Ground Anything Anywhere at Any Granularity
Open source demo platform where you can easily showcase your AI models
Autoregressive Model Beats Diffusion
A Pioneering Open-Source Alternative to GPT-4o
Chat & pretrained large vision language model
An open-source framework for training large multimodal models