VGGSfM: Visual Geometry Grounded Deep Structure From Motion
Language modeling in a sentence representation space
Diversity-driven optimization and large-model reasoning ability
High-Fidelity and Controllable Generation of Textured 3D Assets
Multi-modal large language model designed for audio understanding
Large Multimodal Models for Video Understanding and Editing
LLM-based Reinforcement Learning audio edit model
Easy Docker setup for Stable Diffusion with user-friendly UI
ChatGPT interface with better UI
AI-powered tool to quickly remove watermarks from images flawlessly
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation
Powerful open source image generation model
The ChatGPT Retrieval Plugin lets you easily find personal documents
Stable Diffusion with Core ML on Apple Silicon
Towards Real-World Vision-Language Understanding
Official DeiT repository
Chinese LLaMA-2 & Alpaca-2 Large Model Phase II Project
Example Discord bot written in Python that uses the completions API
Official code for Style Aligned Image Generation via Shared Attention
Code for the paper Hybrid Spectrogram and Waveform Source Separation
Fine-tuning ChatGLM-6B with PEFT
Chinese LLaMA & Alpaca large language model + local CPU/GPU training
800,000 step-level correctness labels on LLM solutions to MATH problem
PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)
A method to increase the speed and lower the memory footprint