Making RAG Simpler with Small and Open-Sourced Language Models
Marrying Grounding DINO with Segment Anything & Stable Diffusion
End-to-end pipeline converting generative videos
A tool to use the Ai2 Open Coding Agents Soft-Verified Agents
Browser automation for AI agents and humans
Persistent context and multi-instance coordination
Block Diffusion for Ultra-Fast Speculative Decoding
Multimodal embedding and reranking models built on Qwen3-VL
SimpleMem: Efficient Lifelong Memory for LLM Agents
A New Axis of Sparsity for Large Language Models
The knowledge and task management backbone for AI coding assistants
Open-source infrastructure for Computer-Use Agents. Sandboxes
"Big Model" trains a visual multimodal VLM with 26M parameters
Simplifies the local serving of AI models from any source
Collection of Gemma 3 variants that are trained for performance
Collection of reference environments, offline reinforcement learning
LLM training in simple, raw C/CUDA
A simple, secure MCP-to-OpenAPI proxy server
Implementation of "MobileCLIP" CVPR 2024
Code release for Cut and Learn for Unsupervised Object Detection
Official implementation of Watermark Anything with Localized Messages
High-resolution models for human tasks
Code for the paper "Evaluating Large Language Models Trained on Code"
Tool for exploring and debugging transformer model behaviors
CLIP, Predict the most relevant text snippet given an image