tiktoken is a fast BPE tokeniser for use with OpenAI's models
This repository contains the official implementation of FastVLM
Official repository for LTX-Video
Hackable and optimized Transformers building blocks
Unified Multimodal Understanding and Generation Models
Qwen2.5-VL is the multimodal large language model series
Audio Language Models are Few-Shot Learners
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
Multi-modal large language model designed for audio understanding
State-of-the-art (SoTA) text-to-video pre-trained model
Large-language-model & vision-language-model based on Linear Attention
Chinese LLaMA & Alpaca large language model + local CPU/GPU training
A minimal PyTorch re-implementation of the OpenAI GPT