CLIP ViT-bigG/14: Zero-shot image-text model trained on LAION-2B
Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video
Fast 12B image model for high-quality text-to-image generation
Multimodal Transformer for document image understanding and layout
Unified multimodal Gemma model for local coding and reasoning
CLIP model fine-tuned for zero-shot fashion product classification
Lightweight multimodal translation model for 55 languages
Text-to-image model optimized for artistic quality and safe generation
Compact agentic model for coding, tools, and productivity tasks
Large agentic model for coding, tools, research, and execution
Open code agent for Lean 4 proofs and formal software verification
Open code agent for Lean 4 proofs and formal software verification
Google’s flagship dense multimodal model for coding and reasoning
4-bit Command A+ model for enterprise agents and multilingual tasks
Open, non-commercial SDXL model for quality image generation
themaCreator - create posts from files
Dense multimodal Qwen model for coding, agents, and long context
Efficient 320B multimodal MoE model for coding and autonomous agents
Local multimodal 30B model for autonomous agents, coding, and tools
FP8 Qwen model for efficient multimodal coding and agent tasks
Versatile 8B-base multimodal LLM, flexible foundation for custom AI
Base Krea image model for LoRA training and fine-tuning
Multimodal 7B model for image, video, and text understanding tasks