Native and Compact Structured Latents for 3D Generation
This repo contains the code for 1D tokenizer and generator
High-Resolution 3D Assets Generation with Large Scale Diffusion Models
Turn your PC, Mac, or Linux box into an AI server.
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Edit Banana: A framework for converting statistical figures
Lets make video diffusion practical
Mixture-of-Experts Vision-Language Models for Advanced Multimodal
AI video customer acquisition and AI short drama creation platform
AI generative media user experience highlighting use of APIs
Multimodal-Driven Architecture for Customized Video Generation
Generating Immersive, Explorable, and Interactive 3D Worlds
One-ink editorial print image skill
[CVPR 2026 Oral] VGGT Omega
CogView4, CogView3-Plus and CogView3(ECCV 2024)
Make drawing and labeling bounding boxes easy as cake
Official MiniMax Model Context Protocol (MCP) server
HivisionIDPhotos: a lightweight and efficient AI ID photos tools
GPT4V-level open-source multi-modal model based on Llama3-8B
Effortless data labeling with AI support from Segment Anything
Essential nodes that are weirdly missing from ComfyUI core
Kaggle Python docker image
Collection of Gemma 3 variants that are trained for performance
Reference PyTorch implementation and models for DINOv3
Automatically find issues in image datasets