Official repository for LTX-Video
Native and Compact Structured Latents for 3D Generation
Implementation of "MobileCLIP" CVPR 2024
Python inference and LoRA trainer package for the LTX-2 audio–video
Python SDK for Claude Agent
AI cognitive-enhancement Skills based on Anthropic's J-space
PyTorch code and models for the DINOv2 self-supervised learning
Unified Multimodal Understanding and Generation Models
26m function call model that runs on incredibly small devices
Multimodal embedding and reranking models built on Qwen3-VL
Generate Any 3D Scene in Seconds
Foundation Models for Time Series
tiktoken is a fast BPE tokeniser for use with OpenAI's models
GLM-4.6V/4.5V/4.1V-Thinking, towards versatile multimodal reasoning
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
RGBD video generation model conditioned on camera input
Large-language-model & vision-language-model based on Linear Attention
ChatGPT interface with better UI
Towards Real-World Vision-Language Understanding
Real-time behaviour synthesis with MuJoCo, using Predictive Control
A minimal PyTorch re-implementation of the OpenAI GPT
Code release for "Masked-attention Mask Transformer
A mix of GAN implementations including progressive growing
Dual LSTM Encoder for Dialog Response Generation