GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Image generation model with single-stream diffusion transformer
Flux 2 image generation model pure C inference
LTX-Video Support for ComfyUI
Visual Causal Flow
Advancing Open-source World Models
Diversity-driven optimization and large-model reasoning ability
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding
Moonshot's most powerful AI model
An experimental version of DeepSeek model
Clean and efficient FP8 GEMM kernels with fine-grained scaling
Infinite Worlds with Versatile Interactions
Codex plugin that turns attached object images into code-only
Ling is a MoE LLM provided and open-sourced by InclusionAI
Open-source image generative foundation model
Qwen3-VL, the multimodal large language model series by Alibaba Cloud
Plugin and skin collection for DeepSeek Harness (DSH) Web UI
CLIP, Predict the most relevant text snippet given an image
Recovering the Visual Space from Any Views
4M: Massively Multimodal Masked Modeling
One-click local MCP server installation in desktop apps
Claude Code image, a one-stop open source transit service
Accurate × Fast × Comprehensive
Designed for text embedding and ranking tasks
Large Multimodal Models for Video Understanding and Editing