Spark-TTS Inference Code
Analyzing Hacker News discussions from a decade ago in hindsight
Fast-stable-diffusion + DreamBooth
Ultimate meta-skill for generating best-in-class Claude Code skills
End-to-end pipeline converting generative videos
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Motion-controllable Video Generation via Latent Trajectory Guidance
A tool to use the Ai2 Open Coding Agents Soft-Verified Agents
Multimodal embedding and reranking models built on Qwen3-VL
A New Axis of Sparsity for Large Language Models
Anthropic's original performance take-home, now open for you to try
Open-source infrastructure for Computer-Use Agents. Sandboxes
"Big Model" trains a visual multimodal VLM with 26M parameters
Improve human sleep through scientifically
Collection of reference environments, offline reinforcement learning
LLM training in simple, raw C/CUDA
Fast and accurate AI powered file content types detection
Implementation of "MobileCLIP" CVPR 2024
Official implementation of Watermark Anything with Localized Messages
High-resolution models for human tasks
CLIP, Predict the most relevant text snippet given an image
Ling is a MoE LLM provided and open-sourced by InclusionAI
Multimodal-Driven Architecture for Customized Video Generation
Personalize Any Characters with a Scalable Diffusion Transformer
Talk to Your AI Agents from Anywhere