Collection of Gemma 3 variants that are trained for performance
Text and image to video generation: CogVideoX and CogVideo
Code for running inference with the SAM 3D Body Model 3DB
This repo contains the code for 1D tokenizer and generator
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Native and Compact Structured Latents for 3D Generation
Reference PyTorch implementation and models for DINOv3
AI generative media user experience highlighting use of APIs
Multimodal-Driven Architecture for Customized Video Generation
A lightweight vision library for performing large object detection
Generating Immersive, Explorable, and Interactive 3D Worlds
State-of-the-art diffusion models for image and audio generation
AI PPT Track Terminator, the strongest PPT Skill ever
Lets make video diffusion practical
Tensor search for humans
Official MiniMax Model Context Protocol (MCP) server
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Personalize Any Characters with a Scalable Diffusion Transformer
Train machine learning models within Docker containers
AutoGluon: AutoML for Image, Text, and Tabular Data
Kaggle Python docker image
Implementation of Imagen, Google's Text-to-Image Neural Network
Easily compute clip embeddings and build a clip retrieval system
Fast-stable-diffusion + DreamBooth
Official implementation of Watermark Anything with Localized Messages