Qwen-Image is a powerful image generation foundation model
Qwen3-omni is a natively end-to-end, omni-modal LLM
Capable of understanding text, audio, vision, video
Guiding Instruction-based Image Editing via Multimodal Large Language
CogView4, CogView3-Plus and CogView3(ECCV 2024)
AI-powered code assistant for Vim. OpenAI and ChatGPT plugin for Vim
Code and models for ICML 2024 paper, NExT-GPT
Tensor search for humans
GPT4V-level open-source multi-modal model based on Llama3-8B
Designed for text embedding and ranking tasks
Gemma open-weight LLM library, from Google DeepMind
GLM-4-Voice | End-to-End Chinese-English Conversational Model
The Multi-Agent Framework
Multilingual sentence & image embeddings with BERT
A high-quality PDF to Markdown tool based on large language model
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System
Qwen2.5-VL is the multimodal large language model series
95% token savings. 155x faster queries. 16 languages
Phi-3.5 for Mac: Locally-run Vision and Language Models
Large-language-model & vision-language-model based on Linear Attention
Repo of Qwen2-Audio chat & pretrained large audio language model
Open source libraries and APIs to build custom preprocessing pipelines
Central interface to connect your LLM's with external data
Practical productivity tools for Claude Code, Codex-CLI
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning