Open-source industrial-grade ASR models
Ultimate meta-skill for generating best-in-class Claude Code skills
Motion-controllable Video Generation via Latent Trajectory Guidance
SimpleMem: Efficient Lifelong Memory for LLM Agents
Improve human sleep through scientifically
Less Code, Lower Barrier, Faster Deployment
Official implementation of Watermark Anything with Localized Messages
Video understanding codebase from FAIR for reproducing video models
Personalize Any Characters with a Scalable Diffusion Transformer
Generic templated configuration management for Kubernetes
Modular quant framework
Open-source AI marketing skills for Claude Code
I Agent designed to interact with ROS1- and ROS2-based robotics system
A personal context-agent that learns how you work
Controllable and fast Text-to-Speech for over 7000 languages
Unified Multimodal Understanding and Generation Models
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Educational framework exploring multi-agent orchestration
Generating Immersive, Explorable, and Interactive 3D Worlds
CTFs as you need them
Framework for managing and maintaining multi-language pre-commit hooks
The lightweight PyTorch wrapper for high-performance AI research
CLI tool to build, test, debug, and deploy Serverless applications
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning