Official repository for LTX-Video
Implementation of "MobileCLIP" CVPR 2024
RGBD video generation model conditioned on camera input
Python inference and LoRA trainer package for the LTX-2 audio–video
Python SDK for Claude Agent
Unified Multimodal Understanding and Generation Models
Foundation Models for Time Series
PyTorch code and models for the DINOv2 self-supervised learning
Generate Any 3D Scene in Seconds
Towards Real-World Vision-Language Understanding
Instructions on how to use the Realtime API on Microcontrollers
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Large-language-model & vision-language-model based on Linear Attention
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
Code release for "Masked-attention Mask Transformer