Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Effortless data labeling with AI support from Segment Anything
Diffusion Transformer with Fine-Grained Chinese Understanding
Gracefully face hCaptcha challenge with multimodal llms
Generate high-definition story short videos with one click using AI
GPT4V-level open-source multi-modal model based on Llama3-8B
Focus on prompting and generating
Codex plugin that turns attached object images into code-only
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
AI Toolkit for Healthcare Imaging
Offline inference engine for art, real-time voice conversations
Nexa SDK is a comprehensive toolkit for supporting ONNX and GGML
Dealing with all unstructured data, such as reverse image search
A Unified Framework for Image Customization
Capable of understanding text, audio, vision, video
Sharp Monocular Metric Depth in Less Than a Second
High-Resolution Image Synthesis with Latent Diffusion Models
Codex plugin that turns attached object images into code-only
Turn your PC, Mac, or Linux box into an AI server.
Recovering the Visual Space from Any Views
Aider is AI pair programming in your terminal
Implementation of 'lightweight' GAN, proposed in ICLR 2021
Official Python inference and LoRA trainer package
Contexts Optical Compression
A general fine-tuning kit geared toward image/video/audio diffusion