This repo contains the code for 1D tokenizer and generator
Build Vision Agents quickly with any model or video provider
Reading book source
Python library and CLI tool to interface with Google Translate
Plug-and-play library to enable agents to call MCP and UTCP tools
Diversity-driven optimization and large-model reasoning ability
Deploy and share agents with open infrastructure
An MCP server that autonomously evaluates web applications
Helping you get the most out of AWS, wherever you use MCP
A solution to build and deploy MCP agents and applications
Multi-Modal Neural Networks for Semantic Search, based on Mid-Fusion
Open source framework for deep learning satellite and aerial imagery
DeepMind's software stack for physics-based simulation
Toolkit for conversational AI
Build cross-modal and multimodal applications on the cloud
A Python toolbox for scalable outlier detection
Stanford NLP Python library for many human languages
High-Fidelity and Controllable Generation of Textured 3D Assets
Multi-modal large language model designed for audio understanding
Large Multimodal Models for Video Understanding and Editing
A minimal yet professional single agent demo project
ComfyUI integration for Microsoft's VibeVoice text-to-speech model
On-device Speech-to-Intent engine powered by deep learning
Swirl queries any number of data sources with APIs
LLM-based Reinforcement Learning audio edit model