State-of-the-art Image & Video CLIP, Multimodal Large Language Models
GPT4V-level open-source multi-modal model based on Llama3-8B
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning
Renderer for the harmony response format to be used with gpt-oss
Qwen2.5-VL is the multimodal large language model series
Real-time Claude Code usage monitor with predictions and warnings
Agent toolkit providing semantic retrieval and editing capabilities
Official SeedVR2 Video Upscaler for ComfyUI
Generates original ARC-AGI-1-style tasks distribution-matched
Codex plugin that turns attached object images into code-only
Audio Language Models are Few-Shot Learners
Build your autonomous hedge fund in minutes
Any model. Any hardware. Zero compromise
Open source RAG framework for building scalable modular AI apps
Sample applications for Google Kubernetes Engine (GKE)
Agent-ready RPA suite with visual workflow automation tools engine
SpikingJelly is an open-source deep learning framework
A comprehensive quantitative trading system with AI-powered analysis
Open-source industrial-grade ASR models
Codebase to Tutorial
Fast-stable-diffusion + DreamBooth
Ultimate meta-skill for generating best-in-class Claude Code skills
End-to-end pipeline converting generative videos
Motion-controllable Video Generation via Latent Trajectory Guidance
Multimodal embedding and reranking models built on Qwen3-VL