Redundancy-aware KV Cache Compression for Reasoning Models
Block Diffusion for Ultra-Fast Speculative Decoding
A high-throughput and memory-efficient inference and serving engine
Achieving 3+ generation speedup on reasoning tasks
Omnilingual ASR Open-Source Multilingual SpeechRecognition
A unified library of SOTA model optimization techniques
The RF and reverse engineering framework for everyone
High-performance Inference and Deployment Toolkit for LLMs and VLMs
Document Image Parsing via Heterogeneous Anchor Prompting”
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Official repository for LTX-Video
State-of-the-art Image & Video CLIP, Multimodal Large Language Models
SGLang is a fast serving framework for large language models
Data manipulation and transformation for audio signal processing
GLM-4.5: Open-source LLM for intelligent agents by Z.ai
Audio Language Models are Few-Shot Learners
TensorRT LLM provides users with an easy-to-use Python API
Official implementation of Watermark Anything with Localized Messages
Scalable data pre processing and curation toolkit for LLMs
kaldi-asr/kaldi is the official location of the Kaldi project
Ultra-Efficient LLMs on End Device
SOTA discrete acoustic codec models with 40/75 tokens per second
Chinese Llama-3 LLMs) developed from Meta Llama 3
A subtitle generator for Japanese Adult Videos.
Code for the paper Language Models are Unsupervised Multitask Learners