PyTorch code and models for the DINOv2 self-supervised learning
LTX-Video Support for ComfyUI
Flux 2 image generation model pure C inference
Image generation model with single-stream diffusion transformer
CLIP, Predict the most relevant text snippet given an image
Diversity-driven optimization and large-model reasoning ability
Large Multimodal Models for Video Understanding and Editing
Proxy that exposes Antigravity provided claude / gemini models
Accurate × Fast × Comprehensive
Visual Causal Flow
One-click local MCP server installation in desktop apps
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
A 0.1B Omni model trained from scratch
Access to Anthropic's safety-first language model APIs
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Recovering the Visual Space from Any Views
Ling is a MoE LLM provided and open-sourced by InclusionAI
An experimental version of DeepSeek model
Claude Code image, a one-stop open source transit service
State of the art LLM and coding model
4M: Massively Multimodal Masked Modeling
Qwen3-VL, the multimodal large language model series by Alibaba Cloud
A Powerful Native Multimodal Model for Image Generation
Clean and efficient FP8 GEMM kernels with fine-grained scaling
Designed for text embedding and ranking tasks