High-Resolution 3D Assets Generation with Large Scale Diffusion Models
AI tool that removes hardcoded subtitles and text from videos locally
Official inference repo for FLUX.2 models
Official Python inference and LoRA trainer package
Foundational video generation model with 13.6B parameters
Recovering the Visual Space from Any Views
MiniMax H3 is a general-purpose, omni-modal generative system
High-Resolution Image Synthesis with Latent Diffusion Models
Reverse engineering Gemini's SynthID detection
This repository contains the official implementation of FastVLM
Temporal-Consistent Diffusion Model for Real-World Video
Repo for SeedVR2 & SeedVR
OCRmyPDF adds an OCR text layer to scanned PDF files
Synthesizing and manipulating 2048x1024 images with conditional GANs
Qwen2.5-VL is the multimodal large language model series
A Customizable Image-to-Video Model based on HunyuanVideo
Generate high-definition story short videos with one click using AI
Knowledge Graph Generation from Any Text
Wan2.1: Open and Advanced Large-Scale Video Generative Model
Native and Compact Structured Latents for 3D Generation
Stable Diffusion web UI
A full spaCy pipeline and models for scientific/biomedical documents
Official repository for LTX-Video
Make any agent harness multimodal-native
GPT4V-level open-source multi-modal model based on Llama3-8B