Motion-controllable Video Generation via Latent Trajectory Guidance
Block Diffusion for Ultra-Fast Speculative Decoding
Multimodal embedding and reranking models built on Qwen3-VL
Minimal Claude Code alternative. Single Python file, zero dependencies
Anthropic's original performance take-home, now open for you to try
Open-source infrastructure for Computer-Use Agents. Sandboxes
"Big Model" trains a visual multimodal VLM with 26M parameters
Collection of reference environments, offline reinforcement learning
Simple and easily configurable grid world environments
Fast and accurate AI powered file content types detection
Implementation of "MobileCLIP" CVPR 2024
VMZ: Model Zoo for Video Modeling
Official implementation of Watermark Anything with Localized Messages
Training Large Language Model to Reason in a Continuous Latent Space
Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
Multimodal-Driven Architecture for Customized Video Generation
Personalize Any Characters with a Scalable Diffusion Transformer
Talk to Your AI Agents from Anywhere
Scalable data pre processing and curation toolkit for LLMs
Build MLOps Pipelines in Minutes
Gorilla: An API store for LLMs
A toolkit to optimize ML models for deployment for Keras & TensorFlow