Movie metadata scraper and organizer for media libraries and NFO
Ready-to-run Docker images containing Jupyter applications
A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
Factorio headless server in a Docker container
Docker image used to run data processing workloads
Automating packaging and software distribution on macOS
A file based wiki that uses markdown
Gemma open-weight LLM library, from Google DeepMind
Generate pixel-perfect macOS folder icons in the native style
Pretrained model hub for Keras 3
This is a background removing tool powered by InSPyReNet
Implementation of Vision Transformer, a simple way to achieve SOTA
Minimal scripts to run the emulator in a container for various systems
PyTorch code and models for V-JEPA self-supervised learning from video
We write your reusable computer vision tools
Implementation of Denoising Diffusion Probabilistic Model in Pytorch
Make any agent harness multimodal-native
Open-source evaluation toolkit of large multi-modality models (LMMs)
PaddlePaddle End-to-End Development Toolkit
Advancing Open-source World Models
Bring the notion of Model-as-a-Service to life
Multimodal AI chat app with dynamic conversation routing
Open source libraries and APIs to build custom preprocessing pipelines