Parse files for optimal RAG
Stable Diffusion web UI
A Multi-Modal World Model for Reconstructing, Generating, Simulation
Official MiniMax Model Context Protocol (MCP) server
Open-source image generative foundation model
Guiding Instruction-based Image Editing via Multimodal Large Language
An open-source toolkit for monitoring Language Learning Models (LLMs)
Collection of Gemma 3 variants that are trained for performance
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Implementation of Imagen, Google's Text-to-Image Neural Network
Text and image to video generation: CogVideoX and CogVideo
A nearly-live implementation of OpenAI's Whisper
Ready-to-use OCR with 80+ supported languages
Free, high-quality text-to-speech API endpoint to replace OpenAI
Framework for building neural networks
Stable Diffusion web UI
Easily compute clip embeddings and build a clip retrieval system
Generating Immersive, Explorable, and Interactive 3D Worlds
Contexts Optical Compression
State-of-the-art (SoTA) text-to-video pre-trained model
AI PPT Track Terminator, the strongest PPT Skill ever
Machine learning, conversational dialog engine for creating chat bots
Flexible Photo Recrafting While Preserving Your Identity
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences