Qwen-Image is a powerful image generation foundation model
All-in-one WebUI for AI generative image and video creation
Guiding Instruction-based Image Editing via Multimodal Large Language
Flutter-based cross-platform app integrating major AI models
CogView4, CogView3-Plus and CogView3(ECCV 2024)
GPT4V-level open-source multi-modal model based on Llama3-8B
The free, Open Source alternative to OpenAI, Claude and others
AI-powered code assistant for Vim. OpenAI and ChatGPT plugin for Vim
Tensor search for humans
Full stack framework for building cross-platform mobile AI apps
Gracefully face hCaptcha challenge with multimodal llms
Capable of understanding text, audio, vision, video
An LLM-based presentation generation platform
Phi-3.5 for Mac: Locally-run Vision and Language Models
Fast and efficient unstructured data extraction
Open source libraries and APIs to build custom preprocessing pipelines
Production-ready AI chat. Start here and make it your own
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System
Moonshot's most powerful AI model
Multilingual sentence & image embeddings with BERT
The open source codebase powering HuggingChat
The Multi-Agent Framework
Qwen3-omni is a natively end-to-end, omni-modal LLM
Skywork-R1V is an advanced multimodal AI model series
A powerful tool for creating datasets for LLM fine-tuning