Python bindings for llama.cpp
Official repository for LTX-Video
Recovering the Visual Space from Any Views
LTX-Video Support for ComfyUI
Run Bonsai (1-bit) and Ternary-Bonsai language models locally
Video Object and Interaction Deletion
Contexts Optical Compression
Bidirectional token-classification model for identifiable info
Visual Causal Flow
Implementation of the Surya Foundation Model for Heliophysics
Open-source multi-speaker long-form text-to-speech model
A theoretical reconstruction of the Claude Mythos architecture
The official repo of Qwen chat & pretrained large language model
Ultra-Efficient LLMs on End Device
Sharp Monocular Metric Depth in Less Than a Second
Use ChatGPT to summarize the arXiv papers
Video understanding codebase from FAIR for reproducing video models
Audio foundation model excelling in audio understanding
Official implementation of DreamCraft3D
Open Source Speech Language Model
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Diffusion Transformer with Fine-Grained Chinese Understanding
Large Multimodal Models for Video Understanding and Editing
Large-language-model & vision-language-model based on Linear Attention
Di♪♪Rhythm: Blazingly Fast & Simple End-to-End Song Generation