Designing Data-Intensive Application
Qwen2.5-VL is the multimodal large language model series
A collection of cv and resume templates written in LaTeX
Bibtex parser for Python 3
Calculate quality metrics with FFmpeg (SSIM, PSNR, VMAF, VIF)
Reverse engineering Gemini's SynthID detection
A python tool for downloading manga from Toonily
A python tool that uses GPT-4, FFmpeg, and OpenCV
Pure Python FFmpeg-based live video / audio streaming to YouTube
ComfyUI wrapper nodes for HunyuanVideo
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
simplejson is a simple, fast, extensible JSON encoder/decoder
Audio Language Models are Few-Shot Learners
Implementation of Vision Transformer, a simple way to achieve SOTA
Hackable and optimized Transformers building blocks
State-of-the-art (SoTA) text-to-video pre-trained model
Unifying 3D Mesh Generation with Language Models
SOTA discrete acoustic codec models with 40/75 tokens per second
Unified Multimodal Understanding and Generation Models
Multi-modal large language model designed for audio understanding
Large-language-model & vision-language-model based on Linear Attention
A high quality MP3 encoder
CGRU: Afanasy render farm manager and RULES project tracker.
XSSer: Cross Site Scripter