Python bindings for llama.cpp
Official repository for LTX-Video
Recovering the Visual Space from Any Views
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
LTX-Video Support for ComfyUI
Open sidebar foundation, supports third-party extensions
From Vibe Coding to Agentic Engineering
Video Object and Interaction Deletion
Contexts Optical Compression
Bidirectional token-classification model for identifiable info
Implementation of the Surya Foundation Model for Heliophysics
Visual Causal Flow
Open-source multi-speaker long-form text-to-speech model
A multimodal model for brain response prediction
A theoretical reconstruction of the Claude Mythos architecture
Sharp Monocular Metric Depth in Less Than a Second
Ultra-Efficient LLMs on End Device
Use ChatGPT to summarize the arXiv papers
Video understanding codebase from FAIR for reproducing video models
Audio foundation model excelling in audio understanding
Multimodal model achieving SOTA performance
Official implementation of DreamCraft3D
Your clothes, extracted and organized with gpt-image
Open Source Speech Language Model
Diffusion Transformer with Fine-Grained Chinese Understanding