⚡ Building applications with LLMs through composability ⚡
Qwen3-Coder is the code version of Qwen3
Qwen-Image is a powerful image generation foundation model
Redundancy-aware KV Cache Compression for Reasoning Models
How to optimize some algorithm in cuda
Open-source evaluation toolkit of large multi-modality models (LMMs)
Designed for text embedding and ranking tasks
File Parser optimised for LLM Ingestion with no loss
Multimodal Agents as Smartphone Users, an LLM-based multimodal agent
Operating LLMs in production
LLM inference server with continuous batching & SSD caching
Generative AI reference workflows
GPT4V-level open-source multi-modal model based on Llama3-8B
MobileLLM Optimizing Sub-billion Parameter Language Models
From Paper to Presentation in One Click
the terminal client for Ollama
A python module to repair invalid JSON from LLMs
Unified framework for building enterprise RAG pipelines
Gorilla: An API store for LLMs
Real-time multi-AI collaboration: Claude, Codex & Gemini
GLM-4-Voice | End-to-End Chinese-English Conversational Model
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
A series of math-specific large language models of our Qwen2 series
Claude + Obsidian knowledge companion
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs