Python bindings for llama.cpp
Instant AI Face Swap
Official repository for LTX-Video
LTX-Video Support for ComfyUI
Video Object and Interaction Deletion
Contexts Optical Compression
From Vibe Coding to Agentic Engineering
Recovering the Visual Space from Any Views
Bidirectional token-classification model for identifiable info
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Qwen3.8-Flash-Next on any consumer hardware
A theoretical reconstruction of the Claude Mythos architecture
Sharp Monocular Metric Depth in Less Than a Second
The official repo of Qwen chat & pretrained large language model
Audio foundation model excelling in audio understanding
A multimodal model for brain response prediction
Open-source multi-speaker long-form text-to-speech model
Visual Causal Flow
OCR expert VLM powered by Hunyuan's native multimodal architecture
Use ChatGPT to summarize the arXiv papers
Your clothes, extracted and organized with gpt-image
VGGSfM: Visual Geometry Grounded Deep Structure From Motion
Open Source Speech Language Model
Video understanding codebase from FAIR for reproducing video models
Ultra-Efficient LLMs on End Device