Advancing Open-source World Models
A Systematic Framework for Interactive World Modeling
DeepMind model for tracking arbitrary points across videos & robotics
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
High-Fidelity and Controllable Generation of Textured 3D Assets
OCR expert VLM powered by Hunyuan's native multimodal architecture
High-Resolution Image Synthesis with Latent Diffusion Models
AI-powered tool to quickly remove watermarks from images flawlessly
StudioOllamaUI is a local, portable interface for Ollama
Memory-efficient and performant finetuning of Mistral's models
Easy Docker setup for Stable Diffusion with user-friendly UI
ChatGLM-6B: An Open Bilingual Dialogue Language Model
Video+code lecture on building nanoGPT from scratch
ChatGPT interface with better UI
Official DeiT repository
Example Discord bot written in Python that uses the completions API
Chinese LLaMA-2 & Alpaca-2 Large Model Phase II Project
Let us control diffusion models
Towards Robust Blind Face Restoration with Codebook Lookup Transformer
Fine-tuning ChatGLM-6B with PEFT
Chinese LLaMA & Alpaca large language model + local CPU/GPU training
A GUI tool for generating subtitle from videos, generating srt files
800,000 step-level correctness labels on LLM solutions to MATH problem
Instruct-tune LLaMA on consumer hardware