Mixture-of-Experts Vision-Language Models for Advanced Multimodal
A Unified Framework for Text-to-3D and Image-to-3D Generation
Wan2.2: Open and Advanced Large-Scale Video Generative Model
Infinite Canvas Workbench for AI creation integrates AI generation
Official MiniMax Model Context Protocol (MCP) server
Stable Diffusion web UI
Open-source image generative foundation model
Guiding Instruction-based Image Editing via Multimodal Large Language
Image inpainting tool powered by SOTA AI Model
Collection of Gemma 3 variants that are trained for performance
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Implementation of Imagen, Google's Text-to-Image Neural Network
A light-weight Markdown editor based on React
Text and image to video generation: CogVideoX and CogVideo
Multimodal-Driven Architecture for Customized Video Generation
A pandoc LaTeX template to convert markdown files to PDF or LaTeX
JavaScript OCR and text extraction for images and PDFs
Stable Diffusion web UI
Easily compute clip embeddings and build a clip retrieval system
Readest is a modern, feature-rich ebook reader
Open source clipboard management tools for Windows, Macos and Linux
Generating Immersive, Explorable, and Interactive 3D Worlds
Contexts Optical Compression
Capable of understanding text, audio, vision, video
Simple & Free Wiki Software