Collection of Gemma 3 variants that are trained for performance
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Implementation of Imagen, Google's Text-to-Image Neural Network
Text and image to video generation: CogVideoX and CogVideo
Multimodal-Driven Architecture for Customized Video Generation
Ready-to-use OCR with 80+ supported languages
Stable Diffusion web UI
Easily compute clip embeddings and build a clip retrieval system
Generating Immersive, Explorable, and Interactive 3D Worlds
Contexts Optical Compression
Capable of understanding text, audio, vision, video
AI PPT Track Terminator, the strongest PPT Skill ever
Flexible Photo Recrafting While Preserving Your Identity
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Offline inference engine for art, real-time voice conversations
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
AI-powered code assistant for Vim. OpenAI and ChatGPT plugin for Vim
An open source implementation of CLIP
Tensor search for humans
AutoGluon: AutoML for Image, Text, and Tabular Data
ImageBind One Embedding Space to Bind Them All
Diffusion Transformer with Fine-Grained Chinese Understanding
Fast stable diffusion on CPU and AI PC