Python Computer Vision & Video Analytics Framework With Batteries Incl
Versatile 8B-base multimodal LLM, flexible foundation for custom AI
Source Code for UURT Humanoid KidSize robots
Open multimodal model for coding, agents, and long-context tasks
Lightweight 24B agentic coding model with vision and long context
Vision-language-action model for robot control via images and text
CLIP ViT-bigG/14: Zero-shot image-text model trained on LAION-2B
Efficient 309B omnimodal MoE for coding, agents, vision, and audio
1T omnimodal MoE model for coding, agents, and long-horizon reasoning
Small 3B-base multimodal model ideal for custom AI on edge hardware
Compact 3B-param multimodal model for efficient on-device reasoning
Powerful 14B-base multimodal model — flexible base for fine-tuning
Dense multimodal Qwen model for coding, agents, and long context
Qwen2.5-VL-3B-Instruct: Multimodal model for chat, vision & video
Open VLA model for autonomous driving reasoning and planning
Efficient multimodal MoE model for coding, tools, and reasoning
Omnimodal AI model for agents, coding, and long-context tasks
Compact 8B multimodal instruct model optimized for edge deployment
Frontier-scale 675B multimodal base model for custom AI training